arXiv:cs.LG(机器学习,全量分类)· Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser·· 5 小时前AI 评分36
UNM-DPO:面向直接偏好优化的不确定性归一化边际方法
Uncertainty-Normalized Margins for Direct Preference Optimization
AI 导读
研究者提出 UNM-DPO,将偏好强度边际与可学习的提示词尺度结合,并基于异方差 Bradley-Terry 模型设计 AO 与 WR 两种训练目标,其中 WR 给出了提示词尺度可辨识的充要条件。
正文
Abstract:Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.
| Comments: | 36 pages |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.38647 [cs.LG] |
| (or arXiv:2609.38647v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38647 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sadegh Khorasani [view email]
[v1]
Tue, 29 Sep 2026 23:08:57 UTC (81 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org