arXiv:cs.LG· Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy·· 3 小时前AI 评分46
VERSE:面向 Agent Harness 的经执行验证的自进化优化器
VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
AI 导读
VERSE 是一个经执行验证的自进化优化器,让优化器在提交前测试草稿编辑、回放失败并扰动可疑步骤,同时同步改进执行器 harness 与自身的提示词、技能、工具、hooks 和笔记,模型权重保持固定。
正文
Abstract:Harness evolution improves an LLM agent's prompts, tools, and workflow, while the optimizer's own tools and procedures often remain fixed. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects. Two observations guide our design. In a controlled study, optimizer self-evolution fails to improve performance without execution-based verification, but achieves the best result of that study when verification is available. Across five executors, self-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control. Motivated by these findings, we introduce VERSE, a Verified Self-Evolving optimizer for agent harnesses. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held-out SWE-rebench tasks and newer out-of-distribution tasks in five languages. Its best validation-selected harness reaches 42.3% and 37.7% accuracy, respectively, against 39.2% and 29.3% for the strongest baselines. Code is available at this https URL.
| Comments: | 45 pages, 13 figures, 15 tables |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02616 [cs.AI] |
| (or arXiv:2610.02616v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02616 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zekai Wang [view email]
[v1]
Fri, 2 Oct 2026 00:16:30 UTC (740 KB)
来源:arXiv:cs.LG · arxiv.org