跳到正文
原文
METR:Research(网页)·· 13 小时前AI 评分56

METR 发布 OpenAI o1-mini 与 o1-preview 初步能力评估报告

Details about METR’s preliminary evaluation of o1-preview September 12, 2024 We measured the performance of OpenAI's o1-mini and o1-preview models on our autonomy and AI R&D task suites, and found they did not exceed the capabilities of the best existing public model we've evaluated, though we could not confidently upper-bound their capabilities. Read more

AI 导读

METR 于 2024 年 9 月 12 日发布对 OpenAI o1-mini 和 o1-preview 的初步评估:在通用自主任务套件上两者未超过已评估的最佳公开模型 Claude 3.5 Sonnet,但 METR 表示无法在有限的评估时间内自信地给出能力上界。

来源:METR:Research(网页) · metr.org