跳到正文
Elon Musk· @elonmusk · X·· 2 小时前AI 评分39
AI 导读

Grok 4.7 非常擅长处理大型代码仓库。

正文

Grok 4.7 is excellent at dealing with large code repositories

引用X Freeze@XFreeze
Grok 4.7 takes the TOP 3 spots on VulcanBench Frontier v4, outperforming Fable 5.1, Opus 5.5, GPT-6 Astra, GPT-6.1 Sol and other leading frontier models • Grok 4.7 xHigh - 93.15 • Grok 4.7 High - 92.71 • Grok 4.7 Medium - 92.30 VulcanBench Frontier v4 tests whether AI models can handle difficult repository-level software engineering work rather than simply generate isolated code snippets The benchmark contains 23 behavioural-reconstruction tasks where models have to rebuild legacy software to match its actual behaviour, with functional correctness tested through hidden tests alongside security, lint/complexity and code quality And the task results are just as impressive: • Grok 4.7 xHigh - 23/23 passed • Grok 4.7 High - 22/23 passed • Grok 4.7 Medium - 21/23 passed Grok 4.7 took all #1, #2 and #3 spots
在 X 查看被引用的帖子

来源:Elon Musk · x.com