跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Sadegh Khorasani, Petrus Mikkola, Matthias Grossglauser·· 5 小时前AI 评分36

UNM-DPO:面向直接偏好优化的不确定性归一化边际方法

Uncertainty-Normalized Margins for Direct Preference Optimization

AI 导读

研究者提出 UNM-DPO,将偏好强度边际与可学习的提示词尺度结合,并基于异方差 Bradley-Terry 模型设计 AO 与 WR 两种训练目标,其中 WR 给出了提示词尺度可辨识的充要条件。

正文

View PDF HTML (experimental)

Abstract:Direct preference optimization (DPO) models binary preferences through a Bradley-Terry model with a common noise scale, without explicitly accounting for preference strength or prompt-dependent uncertainty from human feedback. We introduce uncertainty-normalized margin DPO (UNM-DPO), which combines strength-dependent margins with a learned prompt scale. Motivated by a heteroskedastic Bradley-Terry model, we develop two training objectives. Both compare the implicit rewards of preferred and rejected responses, derived from response log-probability ratios to a reference policy. Advantage-only (AO) divides this reward difference by the prompt scale before subtracting the margin; whole-residual (WR) subtracts the margin before dividing by the scale. For the WR comparison model, we establish a necessary and sufficient condition under which known margins make the prompt scale identifiable. We introduce a practical procedure for learning the scale. Building on WR, we introduce ULNM-DPO-WR, which normalizes each response's implicit reward by its length. We evaluate our methods against DPO and related baselines on HelpSteer2 and HelpSteer3, using the Skywork reward model as a judge. With Llama-3.1-8B-Instruct, ULNM-DPO-WR achieves tie-adjusted win rates against matched DPO of 68.00% and 65.31% on evaluation panels, with higher mean rewards and shorter responses on average. On AlpacaEval with a GPT-4.1 judge and GPT-4-Turbo reference answers, the same 8B policy achieves a length-controlled win rate of 21.62%, compared with 16.39% for DPO and 15.30% for SimPO. These results demonstrate the potential of combining preference-strength margins, learned prompt scales, and length normalization for policy optimization.
Comments: 36 pages
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38647 [cs.LG]
  (or arXiv:2609.38647v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38647

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sadegh Khorasani [view email]
[v1] Tue, 29 Sep 2026 23:08:57 UTC (81 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org