arXiv:cs.AI· Yang Qu, Yusheng Han, Chengjia Feng, Handan Liu·· 6 小时前AI 评分31
LSC-DPO:学习信号控制的直接偏好优化
LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization
AI 导读
研究者提出 LSC-DPO,通过动态调节逻辑 DPO 损失中的 sigmoid 学习信号来改进直接偏好优化。在 AlpacaEval 2、MT-Bench 和 Anthropic-HH 上,LSC-DPO 持续优于 DPO 及强偏好优化基线。基于不同系数初始化会产生不同瞬态学习信号轨迹的发现,作者推导出信号预算补偿规则,显著降低了跨初始化性能波动。
正文
Abstract:Direct Preference Optimization (DPO) has become a standard reward-model-free approach for aligning language models with preference data. However, as the scaled preference margin grows during training, the logistic DPO loss becomes progressively less sensitive to further changes. We study DPO from a loss-level geometric perspective and identify the sigmoid factor as a learning signal that characterizes the local sensitivity of the objective. Based on this view, we propose Learning-Signal-Controlled Direct Preference Optimization (LSC-DPO), which dynamically regulates the learning signal near a target regime. A log-space analysis establishes conditions for stable tracking of the target learning-signal regime. Experiments on AlpacaEval 2, MT-Bench, and Anthropic-HH show that LSC-DPO consistently improves over DPO and strong preference-optimization baselines. We further find that different coefficient initializations induce distinct transient learning-signal trajectories even when their later signal levels become similar. Based on this observation, we derive a signal-budget compensation rule that adjusts the target learning signal to compensate for these transient differences. The resulting compensation substantially reduces performance variation across coefficient initializations.
| Comments: | 27 pages, 15 figures, 14 tables |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07592 [cs.AI] |
| (or arXiv:2610.07592v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07592 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yang Qu [view email]
[v1]
Tue, 6 Oct 2026 01:32:40 UTC (1,298 KB)
来源:arXiv:cs.AI · arxiv.org