arXiv:cs.AI· Hadi Mohammadi·· 4 小时前AI 评分52
trajectory-judge:只看结果的 LLM 评审在智能体轨迹评测中漏掉了什么
trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories
AI 导读
论文提出评测 LLM 智能体轨迹评审的新方法,通过向正确运行注入故障并对比干净运行的配对区分度,揭示仅看请求和最终回复的 14B 评审模型对回复不变的故障召回 34% 到 76% 实为误报。
正文
Abstract:A direct test of an LLM judge of agent trajectories injects faults into correct runs and reports recall, per fault type or by whether the fault broke the environment outcome (loud) or not (silent). Such recall can credit a judge with detection it does not have; paired discrimination, its flag rate on the faults minus its rate on the clean runs they came from, exposes this. Our testbed, a deterministic support desk with a scripted oracle and a one-step fault injector, labels all 400 trajectories exactly. A 14B judge shown only the request and final reply scores 34% to 76% recall on four fault types that leave the reply unchanged. There its input is the clean run's, so its paired discrimination is zero and that recall is its flag rate on clean runs. Splitting by outcome survival does not fix this: its loud recall of 84% is a paired +0.393 and its silent recall of 45% a paired +0.048, all from the two fault types that change the reply. Told to check each step, the same model flags every fault of those four types and 0 of 100 clean runs (95% CI up to 3.6%). It does not reliably check the reply: of four invented promises it flags one every time and the other three once in 42 faults. Shown every step but asked only about the reply, it still reaches a paired +0.69 on reply-unchanged faults, against +1.00 when told to check each step. We recommend reporting paired discrimination against clean parents, split by whether the fault reaches the judge's input and by outcome survival, and release the testbed, raw verdicts and analysis pipeline.
| Comments: | Accepted at the NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development (poster). Camera-ready version. 22 pages, 5 figures, 14 tables. Code and data: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2609.00038 [cs.CL] |
| (or arXiv:2609.00038v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.00038 arXiv-issued DOI via DataCite |
Submission history
From: Hadi Mohammadi [view email]
[v1]
Sat, 29 Aug 2026 10:14:05 UTC (72 KB)
[v2]
Fri, 2 Oct 2026 08:41:50 UTC (117 KB)
来源:arXiv:cs.AI · arxiv.org