METR:Research(网页)·· 12 小时前AI 评分57
METR 评估 DeepSeek-V3 自主能力:未发现超越现有模型的危险能力,GPQA 成绩非数据污染
DeepSeek-V3 Evaluation Results February 12, 2025 We evaluated DeepSeek-V3 for dangerous autonomous capabilities and found no evidence of dangerous capabilities beyond those of existing models such as Claude 3.5 Sonnet and GPT-4o. We also confirmed that its performance on GPQA is not due to training data contamination. Read more
AI 导读
METR 发布对开源权重模型 DeepSeek-V3 的评估报告(2025年2月12日),未发现其具备超越 Claude 3.5 Sonnet 和 GPT-4o 等现有模型的危险自主能力。
来源:METR:Research(网页) · metr.org