arXiv:cs.AI(全量分类)· Yixuan Li, Yiyun Zhou, Yao Long Teng, Fuchao Yang, Yanchen Deng, Zhiyi Lyu, Xuyu Dong, Feng Chen, Bo An·· 5 小时前AI 评分62
arXiv 论文:模型与 Agent harness 的适配关系随任务而变
Finding the Right Fit: Model-Harness Interactions across Agent Tasks
AI 导读
arXiv 论文《Finding the Right Fit: Model-Harness Interactions across Agent Tasks》评估 66 种配置。
正文
Abstract:Choosing an agent system means choosing both a language model and the harness through which it acts. We ask whether a strong model, harness, or pairing stays strong when the setting changes. We evaluate 66 configurations: four configurable harnesses (OpenHands, DeepSeek Harness, PI, and openJiuwen) paired with five models on TUA-Bench, ALE-CLI, and Terminal-Bench 4, plus the native Codex-GPT and Claude Code-Claude pairings. Model rankings reverse across harnesses. On Terminal-Bench 4, Claude leads GPT by 7.94 points in OpenHands but trails it by 30.16 points in PI. For four of the five models, the best harness changes from one benchmark to another, yet some pairings hold: openJiuwen gives Kimi its highest score on all three benchmarks, by 5.61 to 11.11 points. A model's own vendor harness is not reliably its best, and higher cost does not reliably buy a higher score. On Terminal-Bench 4, GPT scores higher under PI than under DSH at less than a quarter of the cost per task. Matched trajectories suggest why fit varies. Models start almost all repairs themselves, so much depends on whether the harness hands failures back in a form the model can use. GPT does best with PI's lean scaffold, while Kimi, which often issues malformed tool calls, does best in openJiuwen. We argue that the model, the harness, and the task should be evaluated together, and we release the harness adapters, evaluation code, and all 6,204 scored trajectories at this https URL and this https URL.
| Comments: | 19 pages, 9 figures, 6 tables. Code: this https URL. Data: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.00917 [cs.AI] |
| (or arXiv:2610.00917v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00917 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yixuan Li [view email]
[v1]
Thu, 1 Oct 2026 01:49:59 UTC (393 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org