arXiv:cs.AI· Hao Wang, Haocheng Yang, Licheng Pan, Lei Shen, Xiaoxi Li, Yinuo Wang, Zhichao Chen, Yuan Lu, Haoxuan Li, Zhouchen Lin·· 6 小时前AI 评分45
ImplicitRM:从隐式反馈学习无偏奖励模型用于 LLM 对齐
Unbiased Reward Modeling from Implicit Feedback for LLM Alignment
AI 导读
针对奖励模型依赖昂贵显式反馈的问题,研究者提出 ImplicitRM,可直接从点击、复制、跳过等隐式用户反馈中学习无偏奖励模型。该方法用分层模型将训练样本划分为四个潜在组,并推导出理论上无偏的似然最大化目标,以解决隐式反馈缺少确定负样本和选择偏差两大挑战。实验在多种 LLM 主干和基准数据集上验证其能从隐式反馈学到准确奖励模型,并提升下游 RLHF 任务表现,该工作已被 ICML 2026 接收。
正文
Abstract:Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collect and difficult to scale. This work studies implicit reward modeling, learning reward models from implicit user feedback, such as clicks, copies and skips. While scalable and cost-effective, implicit feedback poses two key challenges: It lacks definitive negative samples, which makes standard positive-negative classification methods inapplicable; It suffers from selection bias, where responses have heterogeneous propensities to elicit feedback, which further obscures definitive negative samples. To address these challenges, we propose ImplicitRM, which learns unbiased reward models from implicit feedback. It stratifies training samples into four latent groups using a stratification model and derives a likelihood-maximization objective that is theoretically unbiased, thereby addressing both challenges. Experiments across diverse LLM backbones and benchmark datasets validate that ImplicitRM learns accurate reward models from implicit feedback and improves performance on downstream RLHF tasks.
| Comments: | Accepted by ICML 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Applications (stat.AP) |
| Cite as: | arXiv:2603.23184 [cs.CL] |
| (or arXiv:2603.23184v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2603.23184 arXiv-issued DOI via DataCite |
|
| Journal reference: | ICML 2026 |
Submission history
From: Hao Wang [view email]
[v1]
Tue, 24 Mar 2026 13:32:14 UTC (1,090 KB)
[v2]
Tue, 6 Oct 2026 12:34:07 UTC (447 KB)
来源:arXiv:cs.AI · arxiv.org