arXiv:cs.CL· Yekun Chai, Qiwei Peng, Haoyi Xiong·· 4 小时前AI 评分55
XiangqiBench:用中国象棋闭环评测 LLM Agent 的论文
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
AI 导读
研究者提出 XiangqiBench,一个可执行的 LLM Agent 闭环评测基准,要求模型在 119 个有强制将杀的象棋残局中对抗引擎防守方完成将杀,并记录了 12 个前沿 LLM 的 8,568 条多轮轨迹。
正文
Abstract:Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1\% of Sighted trials, yet only 13.9\% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7\% pass@3 but only 5.9\% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3\% of accepted simulation calls stop on an illegal move, and in 49.3\% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02425 [cs.CL] |
| (or arXiv:2610.02425v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02425 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yekun Chai [view email]
[v1]
Thu, 1 Oct 2026 19:48:08 UTC (1,455 KB)
来源:arXiv:cs.CL · arxiv.org