METR:Research(网页)·· 13 小时前AI 评分56
METR 评测 DeepSeek-R1:未发现超出 Claude 3.5 Sonnet 等现有模型的危险自主能力
DeepSeek-R1 Evaluation Results March 5, 2025 We evaluated DeepSeek-R1 for dangerous autonomous capabilities and found no evidence of dangerous capabilities beyond those of existing models such as Claude 3.5 Sonnet and GPT-4o. Interestingly, we found that it did not do substantially better than DeepSeek-V3 on our autonomy suite. Read more
AI 导读
METR 发布对 DeepSeek-R1 的危险自主能力评测,未发现其具备超出 Claude 3.5 Sonnet、GPT-4o 等现有模型的危险能力。
来源:METR:Research(网页) · metr.org