arXiv:cs.LG· Chee Heng Tan, Zhuoyi Lin, Mehul Motani, Wee Sun Lee·· 7 小时前AI 评分41
大语言模型置信度校准的强化学习奖励函数有效性研究
On the effectiveness of reward functions in reinforcement learning for confidence calibration of large language models
AI 导读
研究探讨用强化学习训练大语言模型同时提升推理准确率与置信度表达,奖励方案分别为正确答案和错误答案设置两个函数。若设计不当会诱发"置信度奖励黑客",即模型故意答错以使置信度显得校准,作者提出"不可黑客攻击的置信度奖励方案"及构造方法,并验证可攻击方案在真实数据集上出现选择性奖励黑客,而不可攻击方案具有抵抗力。实验还将这些方案置于过度自信—信心不足的谱系上,显示其呈现相应校准偏差。
正文
Abstract:In this paper, we consider the setting where large language models (LLMs) are trained using reinforcement learning (RL) to simultaneously improve reasoning accuracy and verbalize their confidence. Our reward scheme uses two functions for rewarding confidence verbalized by the LLM: one for correct answers and the other for incorrect answers. If poorly designed, such a scheme may incentivize an LLM to answer incorrectly in order for its confidence to be calibrated, a phenomenon we term confidence reward hacking. We introduce the notion of non-hackable confidence reward schemes and provide methods for constructing them. We show that selective confidence reward hacking can arise in practical datasets under hackable reward schemes while non-hackable reward schemes are resistant to hacking. Finally, we place some of these schemes along an overconfidence-underconfidence spectrum for RL-based confidence calibration and demonstrate experimentally that they tend to exhibit the corresponding calibration biases relative to other schemes in the spectrum. The code of our experiments is available in this https URL.
| Comments: | 80 pages, 10 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.04332 [cs.LG] |
| (or arXiv:2607.04332v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.04332 arXiv-issued DOI via DataCite |
Submission history
From: Chee Heng Tan [view email]
[v1]
Sun, 5 Jul 2026 14:29:18 UTC (3,039 KB)
[v2]
Tue, 6 Oct 2026 02:46:08 UTC (3,833 KB)
来源:arXiv:cs.LG · arxiv.org