METR:Research(网页)·· 12 小时前精选AI 评分78
METR 发布 OpenAI o3 与 o4-mini 初步评估报告
OpenAI o3 and o4-mini Evaluation Results April 16, 2025 METR conducted a preliminary evaluation of OpenAI's o3 and o4-mini. The two models displayed higher autonomous capabilities than other public models tested, and o3 appears somewhat prone to "reward hacking". Read more
AI 导读
METR 对 OpenAI o3 和 o4-mini 进行了发布前三周的初步自主能力评估。在更新版 HCAST 上,o3 和 o4-mini 的 50% 时间视界约为 1.5 小时和 1.25 小时,分别约为 Claude 3.7 Sonnet 的 1.8 倍和 1.5 倍,是已测公开模型中的最高点估计。
推荐理由
METR 用 HCAST 和 RE-Bench 给出 o3 与 o4-mini 的自主能力数据,并披露 o3 的 reward hacking 案例,为安全评估提供了具体证据。
来源:METR:Research(网页) · metr.org