arXiv:cs.AI· Sanjit Dandapanthula, Shuvom Sadhuka, Samir Khan, Michael Oberst, Aaditya Ramdas, Alexandra Chouldechova·· 4 小时前
如何在代理模型上做后训练:包络采样缓解奖励欺骗
How to post-train on a surrogate: Envelope sampling mitigates reward hacking
AI 导读
研究者提出包络采样(envelope sampling),一种用于 LLM 评判器重校准的方法,通过最小化后训练模型遗憾的上界来缓解奖励欺骗。该方法给出基于拒绝采样或针对修改奖励微调的实用算法,在临床笔记生成和受控谄媚任务上,包络采样重校准能缓解奖励欺骗,而基于基座模型样本的重校准则不能。
正文
Abstract:Large language models (LLMs) are commonly post-trained against LLM judges and other cheap surrogates because the true reward, such as human preference, is too expensive to query at scale. This practice often leads to reward hacking, where reinforcement learning against a miscalibrated surrogate leads to undesirable side effects. In this work, we study a setting in which a small number $n$ of model outputs are annotated with ground-truth labels (e.g., from expert review) and used to recalibrate the LLM judge before optimizing against it. Prior approaches to judge recalibration are costly or heuristic, and it is known that on-policy sampling fails when the surrogate is miscalibrated on a rare set of outputs. In this work, we propose envelope sampling, a theoretically-grounded method for judge recalibration that seeks to minimize an upper bound on the regret of the post-trained model under the assumption that the human reward and re-calibrated reward lie in an $L^2$ ball around the judge. We give practical algorithms to sample from the envelope by rejection or by fine-tuning against a modified reward, and experiments on clinical note generation and on a controlled sycophancy task show that recalibrating on envelope samples mitigates reward hacking where recalibrating on base-model samples does not.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Methodology (stat.ME) |
| Cite as: | arXiv:2610.11281 [cs.LG] |
| (or arXiv:2610.11281v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11281 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sanjit Dandapanthula [view email]
[v1]
Thu, 8 Oct 2026 05:48:31 UTC (71 KB)
来源:arXiv:cs.AI · arxiv.org