METR:Research(网页)·· 13 小时前AI 评分42
METR 对 GPT-4o 自主能力初步评估详情:77 项任务、约半数失败可修复
Details about METR’s preliminary evaluation of GPT-4o August 7, 2024 We measured the performance of GPT-4o given a simple agent scaffolding on 77 tasks across 30 task families testing autonomous capabilities. Read more
AI 导读
METR 用简单 agent scaffolding 在 30 个任务族共 77 项自主能力任务上评测 GPT-4o,表现强于 Claude 3 Sonnet 和 GPT-4-Turbo,略弱于 Claude 3.5 Sonnet,与人类 30 分钟基线相当。
来源:METR:Research(网页) · metr.org