跳到正文
arXiv:cs.CL· Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Zeyuan Chen, Lianhui Qin, Chien-Sheng Wu·· 3 小时前AI 评分40

CoTrace:面向终端智能体的 harness 与模型协同演化数据配方

CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution

AI 导读

CoTrace 是面向终端智能体的 harness 感知数据配方,通过轨迹路由、来源匹配和课程刷新,让模型训练严格基于与运行时匹配的验证轨迹。在 Tmax promotion split 上,它把 Qwen3.5-9B 从 78 个任务提升到 88 个(SFT),在线强化学习变体达到 90 个,且更小的 harness 匹配语料比跨同类 harness 汇聚的大语料更省算力。

正文

Authors:Jixuan Chen, Jiaxin Zhang, Qinyuan Ye, Yada Pruksachatkun, Haoxiang Zhang, Jingming Zhuo, Yifan Zhang, Yutong Dai, Juntao Tan, Xiangyu Peng, Silvio Savarese, Zeyuan Chen, Lianhui Qin, Chien-Sheng Wu

View PDF HTML (experimental)

Abstract:Terminal-agent capability depends jointly on model weights and the runtime harness that formats prompts, binds tools, and handles error recovery. Existing harness-model co-evolution approaches improve both components, yet often treat trajectories produced during harness search as an undifferentiated replay buffer. This practice overlooks that a trajectory's value for model training depends on the harness under which it was generated. To systematically analyze this interface, we establish an alternating co-evolution framework that decouples harness search and policy training through component-wise promotion decisions. Within this framework, we introduce CoTrace, a harness-aware data recipe that explicitly governs trajectory routing, provenance matching, and curriculum refresh. Under CoTrace, recurring execution failures guide harness synthesis, while policy training is strictly conditioned on verified rollouts matched to the adopted runtime for supervised fine-tuning (SFT) or fresh online interactions for reinforcement learning (RL). On the Tmax promotion split, CoTrace advances Qwen3.5-9B from 78 to 88 solved tasks under supervised fine-tuning while an online reinforcement variant reaches 90. Specifically, a compact harness-matched corpus produces steady model gains at substantially lower compute than much larger corpora pooled across sibling harnesses. Furthermore, evaluations on Terminal-Bench 2.1 and SWE-bench Lite show that out-of-distribution transfer depends fundamentally on harness compatibility, where maintaining consistency between training and evaluation runtimes prevents procedural execution breakdowns observed under foreign scaffolds.
Comments: Preprint. 32 pages, 7 figures, 17 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.10426 [cs.CL]
  (or arXiv:2610.10426v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.10426

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jixuan Chen [view email]
[v1] Wed, 7 Oct 2026 17:06:59 UTC (1,619 KB)

来源:arXiv:cs.CL · arxiv.org