arXiv:cs.AI· Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Noah A. Smith, Pradeep Dasigi, Teng Xiao·· 4 小时前AI 评分49
重新审视 LLM 智能体 harness 演化的评估方法
Rethinking the Evaluation of Harness Evolution for Agents
AI 导读
研究指出自动 harness 演化评估存在两大问题:未在匹配反馈与推理预算下对比简单基线,且用同一 benchmark 的验证信号搜索配置、又在同一 benchmark 上报告结果,违反训练测试分离原则。
正文
Authors:Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Noah A. Smith, Pradeep Dasigi, Teng Xiao
Abstract:Harness evolution is an iterative search procedure that repeatedly evaluates and revises candidate harnesses used for LLM agents using task feedback. We revisit the evaluation of such automatic harness evolution procedures and identify two fundamental issues in the protocol. First, prior work does not compare these approaches with simple task-level search baselines under matched feedback and inference budgets. Second, prior work searches for harness configurations using verification signals (e.g., unit test cases) drawn from the same benchmarks on which it reports the final performance of the evolved harnesses, violating the standard separation between training and test data. To address this, we compare automatic harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and evaluate evolved harnesses on held-out tasks to assess generalization. Following prior work, we experiment on Terminal-Bench 2.1 and find that automatic harness evolution fails to outperform simple test-time scaling methods both with and without test cases, and exhibits limited generalization. However, we find that long-horizon games are a promising setting for automatic harness evolution, as they are difficult enough to leave headroom, rely heavily on adaptation to out-of-distribution dynamics, and provide granular feedback by design. In these settings, task-specific harness evolution improves over the search baseline by 80.0% on ARC-AGI-3 and by 10.9% on EdgeBench under matched budgets. Together, these findings highlight the need for matched-budget baselines and held-out evaluation to distinguish genuine harness improvements from benchmark-specific search and overfitting, and point to a more careful characterization of when automatic harness evolution is actually useful. Our code is available at this https URL.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.12227 [cs.AI] |
| (or arXiv:2607.12227v3 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.12227 arXiv-issued DOI via DataCite |
Submission history
From: Yike Wang [view email]
[v1]
Tue, 14 Jul 2026 00:18:42 UTC (113 KB)
[v2]
Thu, 27 Aug 2026 12:13:50 UTC (113 KB)
[v3]
Thu, 1 Oct 2026 21:04:31 UTC (199 KB)
来源:arXiv:cs.AI · arxiv.org