arXiv:cs.CL· Hexuan Deng, Yue Wang, Wenyu Jiang, Cheng Yang, Haolin Yang, Zhaohua Zhang, Chenchen Zhao, Beiduo Chen, Muxi Chen, Sa Zhu, Geyuan Zhu, Jianhuan Zhuo, Qiuyong Xiao, Tianwen Jiang, Jihong Zhang, Xuebo Liu·· 3 小时前
SWE-Journey:面向编程助手长周期多轮交互的更真实评测基准
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
AI 导读
研究者推出 SWE-Journey 基准,用于更真实地评测编程助手的长周期、多轮交互能力。该基准通过弱到强合成流水线自动构建长周期编程任务,并基于真实交互数据挖掘四类用户画像、搭建用户模拟智能体复现真实交互。评测显示,模型在软件架构师画像下通过超 75% 的功能测试,但在非程序员画像下不足 25%。
正文
Authors:Hexuan Deng, Yue Wang, Wenyu Jiang, Cheng Yang, Haolin Yang, Zhaohua Zhang, Chenchen Zhao, Beiduo Chen, Muxi Chen, Sa Zhu, Geyuan Zhu, Jianhuan Zhuo, Qiuyong Xiao, Tianwen Jiang, Jihong Zhang, Xuebo Liu
Abstract:Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.11559 [cs.CL] |
| (or arXiv:2610.11559v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11559 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hexuan Deng [view email]
[v1]
Thu, 8 Oct 2026 09:21:49 UTC (465 KB)
来源:arXiv:cs.CL · arxiv.org