arXiv:cs.CL· Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova·· 4 小时前AI 评分44
奖励模型记住了什么?EMNLP 2026 研究揭示判别式训练的偏置
What do Reward Models Memorize?
AI 导读
研究通过两个人类偏好数据集上的反事实记忆测量,发现判别式训练的奖励模型(RM)会把记忆错误分配给简单、高margin的偏好对,并记住数据集特有的捷径(如模型身份、用户采样策略)。面对未见偏好对时,RM还会过度泛化长度、顺从性等人类偏好的简单启发式相关性。该研究已被 EMNLP 2026 Findings 接收,表明这类 RM 存在偏置,尚无法在依赖上下文的场景中判断回复质量。
正文
Abstract:This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.
| Comments: | Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026 |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.24484 [cs.LG] |
| (or arXiv:2607.24484v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.24484 arXiv-issued DOI via DataCite |
Submission history
From: Ivo Verhoeven [view email]
[v1]
Mon, 27 Jul 2026 14:20:59 UTC (585 KB)
[v2]
Wed, 7 Oct 2026 14:11:39 UTC (810 KB)
来源:arXiv:cs.CL · arxiv.org