arXiv:cs.AI(全量分类)· Ruoming Jin, Xinyu Li, Hao Zhou, Jianfeng Zhu, Ruixin Guo, Feodor Dragan, Lei Xu, Haixun Wang, Yang Zhou·· 5 小时前AI 评分34
GAP-DPO:面向个性化偏好优化的梯度对齐配对选择方法
Gradient-Aligned Pair Selection for Personalized Preference Optimization
AI 导读
研究者提出 GAP-DPO(Geometry-Aligned Preference DPO),一种通过梯度对齐进行效用感知配对选择的迭代算法,用于个性化偏好优化。
正文
Abstract:Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optimization from explicit user utility and can lead to degraded personalization. We formalize personalized preference learning as a geometry-aligned optimization problem by analyzing the first-order interaction between gradients of expected user utility and DPO update directions. Our analysis reveals that, under off-policy sampling, the DPO update transitions from a purely error-corrective signal to a reinforcement-like update when preference margins are directionally aligned with utility gradients. This perspective exposes pair selection as a geometric decision that governs whether preference optimization advances or hinders personalization. Motivated by this insight, we propose GAP-DPO (Geometry-Aligned Preference DPO), an iterative algorithm that performs utility-aware, geometry-aligned pair selection while controlling distribution shift via epoch-wise regeneration. Experiments on personalized text generation benchmarks show that GAP-DPO consistently improves stylistic fidelity, preference alignment, and generation quality compared to standard DPO variants. Together, our results establish gradient alignment as a unifying principle for personalized preference optimization and demonstrate that pair selection is an intrinsic component of the optimization geometry rather than a heuristic preprocessing step.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00061 [cs.AI] |
| (or arXiv:2610.00061v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00061 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinyu Li [view email]
[v1]
Fri, 4 Sep 2026 05:37:29 UTC (67 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org