arXiv:cs.LG· Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang·· 3 小时前AI 评分69
CriticHack:研究视觉奖励模型在机器人策略优化中放大错误目标操作的问题
CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization
AI 导读
arXiv 论文 CriticHack 发现,用学习型视觉奖励模型优化机器人策略时,奖励模型对操作错误对象的执行打分可能与完成任务一样高,优化会放大这类错误目标失败,而奖励和任务成功率同时上升,常规监控信号看似健康。
正文
Abstract:Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success both rise, so the signals a practitioner would normally monitor look healthy. We fine-tune every denoiser parameter of a diffusion policy against Robometer on a drawer task. Starting from a supervised policy with no prior reward exposure, five training runs raise task success by 10.2 percentage points and wrong-object failures by 10.9 points on 512 evaluation seeds, whereas five runs trained on the simulator's task-completion signal raise success without amplifying wrong-object failures (difference 9.2 points, 95% CI 5.6 to 13.0). The amplification recurs from a policy previously optimized against learned rewards, under the policy's native diffusion sampler, at matched distance from the initial policy, and across constrained-policy experiments with two critics and two optimizers. A tilt model explains when it occurs: under KL-regularized optimization, an outcome becomes more frequent whenever its expected reward under the initial policy exceeds the population average. Robometer separates successes from failures well overall (AUROC .81) but scores wrong-object failures slightly above successes (AUROC .37), so optimization raises both. The same model predicts the outcome shifts across 26 constrained settings (Spearman .89), including those in which task success falls, and Robometer's own published success-termination recipe inherits the error. A frozen outcome verifier redirects the same optimization toward the requested task.
| Comments: | 60 pages |
| Subjects: | Robotics (cs.RO); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02527 [cs.RO] |
| (or arXiv:2610.02527v1 [cs.RO] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02527 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiaxuan Luo [view email]
[v1]
Thu, 1 Oct 2026 21:57:25 UTC (1,962 KB)
来源:arXiv:cs.LG · arxiv.org