跳到正文
arXiv:cs.CL· Andreea Dutulescu, Stefan Ruseti, Mihai Masala, Traian Rebedea, Mihai Dascalu·· 4 小时前AI 评分31

GAW-PO:基于梯度对齐 token 权重的偏好优化方法

GAW-PO: Preference Optimization with Gradient-Aligned Token Weights

AI 导读

GAW-PO 是一种针对 DPO 的梯度对齐 token 重加权方法,通过估计被拒回答中每个 token 的梯度是否与偏好更新方向冲突来调整其惩罚权重。在覆盖数学、推理、编程和问答的 11 个基准上,其平均性能比标准 DPO 高 0.97 分,比最强竞争基线高 0.65 分。随着 DPO 正则参数 β 减小,标准 DPO 性能急剧下降,而 GAW-PO 持续提升。

正文

View PDF HTML (experimental)

Abstract:Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressive language models are optimized token by token. As a result, all tokens in a rejected response contribute to the negative training signal, including tokens that may encode behavior that is useful for the preferred response. We introduce GAW-PO, a gradient-aligned token reweighting method for DPO that estimates, for each rejected token, whether penalizing it would interfere with the preferred update directions. Tokens whose gradients are strongly aligned with the preferred behavior receive a weaker negative contribution, while conflicting tokens retain a stronger penalty. Our method achieves the highest average performance among the evaluated preference-optimization methods, improving by 0.97 points over standard DPO and 0.65 points over the strongest competing baseline across 11 benchmarks spanning mathematics, reasoning, coding, and question answering. We further show that gradient-aligned weighting is substantially more robust to aggressive preference optimization: as the DPO regularization parameter $\beta$ decreases, standard DPO degrades sharply, whereas GAW-PO continues to improve. These results suggest that accounting for the interaction between rejected-token updates and preferred behavior provides an effective form of token-level credit assignment for preference optimization.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.01511 [cs.CL]
  (or arXiv:2610.01511v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.01511

arXiv-issued DOI via DataCite

Submission history

From: Andreea Dutulescu [view email]
[v1] Thu, 1 Oct 2026 11:46:28 UTC (469 KB)
[v2] Fri, 2 Oct 2026 08:02:08 UTC (469 KB)

来源:arXiv:cs.CL · arxiv.org