来自芝加哥大学、复旦大学、清华大学、TensorBlock 等机构的研究者发布 AgentBug-Smith,自动将真实 GitHub 缺陷报告转化为可运行测试,构建含 200 个可复现 harness 缺陷的 LIVE-HARNESS-BENCH 基准。
Self-improving AI agents will need to fix their own code.
And this paper from top US+China labs, shows coding agents miss most such bugs but improve with lessons from past fixes.
that real bugs in agent harnesses, can be automatically turned into a growing set of runnable tests.
An agent's own code is everything around the model: tool calls, memory, and prompts. Its bugs depend on live model calls, which makes them hard to recreate and test.
So the researchers built AgentBug-Smith, which turns real GitHub bug reports into runnable tests. The result is a 200-bug benchmark that keeps growing.
The best of 3 coding agents fixed just 9% of those bugs, versus about 40% reported on regular software bugs. A short guide of lessons from past fixes lifted an agent from 1 to 6 correct fixes on 79 unseen bugs.
Before trusting a coding agent with your agent's code, try it on bugs you've already fixed.
来源:Rohan Paul · x.com