arXiv:cs.AI· Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He·· 3 小时前
谁验证验证者?可检查评分器与自我改进智能体的共同演化
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
AI 导读
研究提出让验证者成为演化对象:由聚类失败合成、对小型确定性缺陷检测器组成的可检查表达式,经出生门控,并依据十项锚定参考集与无标签输出共识而非智能体得分进行筛选。在 MBPP+ 上,其留出集一致性比人工编写的种子组合高出 +0.21,且优于其所包含的裸 LLM 评判器。移除锚定守卫会使验证器退化为恒过的空洞评分器,但其训练出的技能依然同样有效,说明下游任务得分无法认证自我演化的验证器。
正文
Abstract:We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
| Comments: | Accepted at the NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.11464 [cs.AI] |
| (or arXiv:2610.11464v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11464 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xing Zhang [view email]
[v1]
Thu, 8 Oct 2026 08:13:53 UTC (176 KB)
来源:arXiv:cs.AI · arxiv.org