跳到正文
arXiv:cs.LG· Oleksandr Cherednichenko, Roman Klypa·· 5 小时前AI 评分35

Suan:在大语言模型中矫正直接偏好安全对齐

Suan: Rectifying Direct Preference Safety Alignment in Large Language Models

AI 导读

研究者提出偏好优化算法 Suan,将优化目标直接建立在梯度层面、绕过标准变分推导,从而获得更可解释且稳健的训练动态。在多种竞争性基线和方法基准上的评测显示,Suan 在实现更优安全对齐的同时完整保留了模型的回答实用性,缓解了开源权重模型后训练中常见的过度拒答与通用质量下降问题。

正文

View PDF HTML (experimental)

Abstract:Integrating robust safety guardrails into Large Language Models (LLMs) is essential for delivering helpful yet harmless responses. While proprietary systems exhibit reliable safety controls, their underlying methodologies and trade-offs remain largely undisclosed. Achieving comparable security in open-weight models remains a persistent challenge, as post-trained variants frequently suffer from over-refusal and degraded general quality. To overcome these drawbacks, we introduce Suan, a novel preference optimization algorithm. Unlike existing methods, we formulate the optimization objective directly at the gradient level, bypassing the standard variational derivation. As a result, we obtain more interpretable and robust training dynamics. Extensive evaluations across a diverse suite of competitive baselines and benchmarks demonstrate that Suan achieves superior safety alignment while fully preserving response utility.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.08634 [cs.LG]
  (or arXiv:2609.08634v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.08634

arXiv-issued DOI via DataCite

Submission history

From: Oleksandr Cherednichenko [view email]
[v1] Tue, 8 Sep 2026 12:04:38 UTC (426 KB)
[v2] Fri, 2 Oct 2026 10:27:05 UTC (426 KB)

来源:arXiv:cs.LG · arxiv.org