Lee Robinson· @leerob · X·· 20 天前AI 评分27
AI 导读
我们刚刚推出了 CursorBench 4.0! 它包含了新的任务,用于评估模型遵循指令的能力、在具有挑战性的项目上长期工作的表现,并且比之前更难(所以所有模型的得分都更低了)。
正文
We just rolled out CursorBench 4.0!
It includes new tasks for how well models follow instructions, work on challenging projects over time, and is more difficult than before (so all models score lower).
来源:Lee Robinson · x.com