跳到正文
原文
METR:Research(网页)·· 12 小时前AI 评分57

METR 发布 Claude 3.7 Sonnet 自主能力评估报告

Claude 3.7 Evaluation Results April 4, 2025 METR conducted a preliminary evaluation of Claude 3.7 Sonnet. While we failed to find significant evidence for a dangerous level of autonomous capabilities, the model displayed impressive AI R&D capabilities on a subset of RE-Bench which provides the model with ground-truth performance information. We believe these capabilities are central to important threat models and should be monitored closely. Read more

AI 导读

METR 对 Claude 3.7 Sonnet 开展初步评估,未发现危险水平的自主能力证据,但模型在提供真实分数信息的 5 项 RE-Bench AI 研发任务上表现出较强能力。

来源:METR:Research(网页) · metr.org