HuggingFace Daily Papers·· 1 天前AI 评分41
VeriFine:通过扩展验证实现具身推理的自我改进
VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning
AI 导读
VeriFine 是一个 agent harness 框架,通过策略、训练课程与评判器的协同演化来扩展验证能力,解决具身推理中固定评判器制约自我改进的问题。该框架包含 Policy Improvement Loop 与 Judge Improvement Loop,后者在验证成为瓶颈时借助人类指导与 coactive calibration 精炼评判器。
正文
Authors:Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding
Abstract:Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.
| Comments: | Project Website: this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.08761 [cs.AI] |
| (or arXiv:2610.08761v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08761 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zewei Zhou [view email]
[v1]
Tue, 6 Oct 2026 17:50:29 UTC (3,423 KB)
来源:HuggingFace Daily Papers · arxiv.org