METR:Research(网页)·· 13 小时前AI 评分57
METR 研究:单元测试评分高估 AI 智能体真实编码表现
Research Update: Algorithmic vs. Holistic Evaluation August 13, 2025 Many AI benchmarks use algorithmic scoring to evaluate how well AI systems perform on some set of tasks. However, AI systems often produce code that scores well but isn't production-ready due to issues with test coverage, formatting, and code quality. This helps explain why AI tools show less productivity improvement than expected despite strong performance on coding benchmarks. Read more
AI 导读
METR 在来自 stdlib-js 和 hypothesis 两个仓库的 18 个真实任务上,用 Inspect ReAct 智能体运行 Claude 3.7 Sonnet,算法评分下成功率为 38%(±19%,95% CI),但人工抽查的 15 个 PR 中无一可直接合并。
来源:METR:Research(网页) · metr.org