Cognition 模型 / Devin 博客(网页)·· 13 小时前AI 评分59
Cognition 推出 FrontierCode 基准,衡量模型能否写出可合并的高质量代码
Introducing FrontierCode06.08.26Today’s coding benchmarks have established that models can write correct code, but the question we should really be asking is: can models actually write good code?
AI 导读
Cognition 发布 FrontierCode 基准,从正确性、测试质量、范围纪律、风格等维度评估模型代码能否达到开源维护者可合并的标准,误报率比 SWE-Bench Pro 低 81%。
来源:Cognition 模型 / Devin 博客(网页) · cognition.com