跳到正文
arXiv:cs.CL· Radhika Gaonkar·· 4 小时前

TRACE:诊断智能体评估中的验证器脆弱性

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

AI 导读

研究团队提出 TRACE 协议,将 LLM 智能体评估中的分数变化从结论转为可检验的诊断,用于区分分数波动源于智能体能力变化还是评估本身。在 25 个合成任务中,仅重命名工具就让脚本化智能体得分下降 0.250,恢复原名后差距完全消失。

正文

View PDF HTML (experimental)

Abstract:Verifier scores now serve as both benchmark metrics and training rewards for large language model (LLM) agents, and a change in score is routinely read as a change in capability. It may instead reflect a change in the evaluation. We introduce TRACE, a protocol that turns a score change from a verdict into a testable diagnosis: it applies a targeted change to one part of an evaluation, compares paired runs, checks whether the agent's behavior changed, and rescores unchanged trajectories to test whether the scoring rule is responsible. In a controlled suite of 25 synthetic tasks, renaming tools lowers a scripted agent's score by 0.250 even though it performs exactly the same operations; restoring the original names at scoring time closes the entire gap, while the same mutation exposes a genuine behavioral failure in a second agent. On public $\tau^2$-bench tasks with four LLM agents, an initial 30-task study finds mixed reward changes whose one clear effect does not replicate. In a larger follow-up on 88 new tasks with repeated runs per condition, renaming tools or reformatting tool outputs leaves reward unchanged to within $\pm$0.10 for seven of eight agent-change pairs, whereas tool names that deliberately mislead lower every agent's reward by 0.20-0.44, showing that the setup can detect real effects. Identical reruns flip 15-36% of task outcomes, so single-run comparisons cannot separate presentation effects from run-to-run variation. Two frontier LLM judges give consistent verdicts when a fixed trajectory is presented differently, yet disagree with each other on 57% of the same records, largely because one grades procedure rather than outcome. TRACE thus separates what a score change says about the agent from what it says about the measurement.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.11678 [cs.CL]
  (or arXiv:2610.11678v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11678

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Radhika Gaonkar [view email]
[v1] Thu, 8 Oct 2026 10:52:26 UTC (157 KB)

来源:arXiv:cs.CL · arxiv.org