arXiv:cs.LG· Xuying Ning, Dongqi Fu, Tianxin Wei, Yuanchen Bei, Xiyuan Yang, Wujiang Xu, Yueqi Song, Bingxuan Li, Zihao Li, Hanqing Zeng, Xiang Shen, Yajuan Wang, Yifan Wu, Qifan Wang, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He·· 7 小时前AI 评分43
EvoHarness-RL:为自进化 Agent 学习运行时 Harness 协调
EvoHarness-RL: Learning Runtime Harness Coordination for Self-Evolving Agents
AI 导读
EvoHarness-RL 是一个将环境相关 harness 实现与统一策略接口分离的框架,通过 Belief、Progress、Experience(BPE)工作区和四个紧凑 harness 动作组织外部支持,并用监督初始化加成本感知 GRPO 让 harness 协调可学习。
正文
Authors:Xuying Ning, Dongqi Fu, Tianxin Wei, Yuanchen Bei, Xiyuan Yang, Wujiang Xu, Yueqi Song, Bingxuan Li, Zihao Li, Hanqing Zeng, Xiang Shen, Yajuan Wang, Yifan Wu, Qifan Wang, Jiayi Liu, Hong Li, Yinglong Xia, Xiangjun Fan, Hanghang Tong, Jingrui He
Abstract:Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, recover from failures, and reuse experience across extended interactions. Yet existing harnesses and their use are often tailored to environments and controlled through prompts, heuristics, or system-specific rules, making agent and harness coordination difficult to jointly optimize. We introduce EvoHarness-RL, a unified framework that separates environment-specific harness implementations from a shared policy-facing interface. EvoHarness-RL organizes external support into a Belief, Progress, and Experience (BPE) workspace and exposes four compact harness actions for accessing and updating this state. We first instantiate BPE as an inference-time scaffold and then make harness coordination learnable through supervised initialization followed by cost-aware GRPO. Across heterogeneous long-horizon tasks, EvoHarness-Base improves the average success rate of frontier models by 10.0 percentage points, while EvoHarness-RL outperforms the strongest open-source baseline by 8.5 percentage points, together with higher RL rollout efficiency and stronger generalization to unseen tasks. Our analyses show that training gradually shifts agents from frequent scaffold use toward selective, environment-dependent harness access as the policy becomes more capable, while the external workspace continues to evolve and refine itself to better support task execution and generalization. Together, these results show that long-horizon agents benefit not only from external scaffolding itself, but also from learning how and when to coordinate with external support as part of a cost-aware policy.
| Comments: | Accepted to LLA@COLM 2026 |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2608.05446 [cs.LG] |
| (or arXiv:2608.05446v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2608.05446 arXiv-issued DOI via DataCite |
Submission history
From: Xuying Ning [view email]
[v1]
Wed, 5 Aug 2026 22:29:20 UTC (1,730 KB)
[v2]
Mon, 5 Oct 2026 22:53:08 UTC (3,191 KB)
来源:arXiv:cs.LG · arxiv.org