跳到正文
arXiv:cs.AI· Xing Zhang, Guanghui Wang, Yanwei Cui, Ziyuan Li, Wei Qiu, Bing Zhu, Peiyang He·· 3 小时前

谁验证验证者?可检查评分器与自我改进智能体的共同演化

Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents

AI 导读

研究提出让验证者成为演化对象:由聚类失败合成、对小型确定性缺陷检测器组成的可检查表达式,经出生门控,并依据十项锚定参考集与无标签输出共识而非智能体得分进行筛选。在 MBPP+ 上,其留出集一致性比人工编写的种子组合高出 +0.21,且优于其所包含的裸 LLM 评判器。移除锚定守卫会使验证器退化为恒过的空洞评分器,但其训练出的技能依然同样有效,说明下游任务得分无法认证自我演化的验证器。

正文

View PDF HTML (experimental)

Abstract:We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and ends ahead of the bare LLM judge it contains. One finding should change how co-evolved verifiers are validated: removing the anchor guards collapses the verifier into a vacuous always-pass grader, yet that collapsed verifier trains skills just as well. Downstream task score cannot certify a self-evolved verifier. Score does answer sufficiency, and there an evolved verifier can substitute: Double Ratchet, pairing the verifier with a lifecycle-managed skill loop, retains 88-110% of the lift that ground truth or a rubric buys the same loop, across code generation, enterprise text-to-SQL, and reference-free report generation. When evolved skills gamed the report rubric, an outer judge caught it and one added detector repaired it; the judge itself was wrong until given the task contract.
Comments: Accepted at the NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.11464 [cs.AI]
  (or arXiv:2610.11464v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11464

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xing Zhang [view email]
[v1] Thu, 8 Oct 2026 08:13:53 UTC (176 KB)

来源:arXiv:cs.AI · arxiv.org