arXiv:cs.AI(全量分类)· Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee·· 5 小时前AI 评分54
VeriHarness:面向长程任务的智能体验证方法
VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks
AI 导读
Google 研究者提出 VeriHarness,通过给 LLM 提供工作区、证据工具和可复用验证技能,将其转化为智能体验证器,用于校验长程任务中多次采样的 rollout 输出。
正文
Abstract:As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts that can contain complementary correct claims, but we need a reliable verification mechanism to determine which claims to trust. We first find that disagreement often exposes correct alternatives, while consensus can conceal errors. These observations motivate VeriHarness, which turns the underlying LLM a generator uses into an agentic verifier by giving it a workspace, evidence tools, and reusable verification skills. A disagreement resolver checks competing claims against environmental evidence, while a consensus challenger tests shared claims and searches for omitted requirements. Their findings guide the selection and revision of the final artifact. Across five long-horizon workspace benchmarks and two frontier models, VeriHarness achieves the highest selection scores among the evaluated baselines. Evidence-backed revision further improves average performance, bringing gains over a single rollout to 6.2 points with Gemini 3.5 Flash and 6.4 points with Claude Opus 4.8. We further show that verification skills can self-improve from failure feedback, demonstrating VeriHarness as a novel and critical approach for scaling long-horizon agentic verification. We release the full pool of approximately 26,000 rollouts across all five benchmarks and both models, produced at a cost of over $100,000, to support future research on agentic verification.
| Subjects: | Artificial Intelligence (cs.AI); Multiagent Systems (cs.MA) |
| Cite as: | arXiv:2610.00972 [cs.AI] |
| (or arXiv:2610.00972v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00972 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rujun Han [view email]
[v1]
Thu, 1 Oct 2026 03:02:17 UTC (630 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org