跳到正文
原文
arXiv:cs.AI(全量分类)· Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou·· 5 小时前AI 评分34

GAP-DPO:面向个性化偏好优化的梯度对齐配对选择方法

Gradient-Aligned Pair Selection for Personalized Preference Optimization

AI 导读

研究者提出 GAP-DPO(Geometry-Aligned Preference DPO),一种通过梯度对齐进行效用感知配对选择的迭代算法,用于个性化偏好优化。

正文

View PDF HTML (experimental)

Abstract:Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00061 [cs.AI]
  (or arXiv:2610.00061v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00061

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinyu Li [view email]
[v1] Fri, 4 Sep 2026 05:37:29 UTC (67 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org