约翰霍普金斯大学与卡内基梅隆大学的论文显示,4B小模型可根据失败报告改写智能体harness代码,并将该技能迁移到新任务。在21个未见推理任务类型上,4B编辑器的平均编辑得分从0.32升至0.62,超过其35B教师模型。在HotpotQA上训练的编辑器还能持续改进另外2个QA基准的harness,且每任务多轮编辑优于单次修复。
New John Hopkins and Carnegie Mellon University Paper Shows that a small model can learn to improve an agent's harness code from run results, and the skill transfers to new tasks.
A small model trained to rewrite an agent's harness code from failure reports can adapt agents to new tasks, so let it tune your harness instead of doing it by hand.
The harness is the code that decides what the model sees and which tools it calls. The editor reads the harness and what failed, then writes a code change, rewarded by how well the new harness scores.
On 21 unseen reasoning task types, a 4B editor's average edit score rose from 0.32 to 0.62, above its 35B teacher. A separate editor trained on HotpotQA kept improving harnesses on 2 other QA benchmarks.
Run it for several rounds per task, since repeated edits beat single fixes.
来源:Rohan Paul · x.com