METR:Research(网页)·· 13 小时前AI 评分42
METR 首份公开报告:评估 LLM 智能体在真实自主任务中的表现
New report: Evaluating Language-Model Agents on Realistic Autonomous Tasks July 31, 2023 We have just released our first public report. It introduces methodology for assessing the capacity of LLM agents to acquire resources, create copies of themselves, and adapt to novel challenges they encounter in the wild. Read more
AI 导读
ARC Evals(现 METR)发布首份公开报告,提出"自主复制与适应"(ARA)评估方法,用 12 项难度递增的真实任务测试基于 Claude 和 GPT-4 的 4 个 LLM 智能体。结果显示这些智能体只能完成最简单的 ARA 任务,在较难任务上仅有部分进展,表明普通用户难以借此造出危险的自主智能体。
来源:METR:Research(网页) · metr.org