arXiv:cs.AI· Jinghao Pang, Jitai Hao, Qiang Huang, Zhaochun Ren, Jun Yu·· 4 小时前AI 评分41
超越拒绝模式:SSRFT 用安全角色内化实现更稳健可泛化的 LLM 安全对齐
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
AI 导读
研究者提出 SSRFT(Supervised Safe-Role Fine-Tuning),首次将 LLM 安全对齐重构为对预设安全角色的内化,而非依赖显式拒绝模式。
正文
Abstract:Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.
| Comments: | 27 pages,7 figures, under review |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2610.07023 [cs.AI] |
| (or arXiv:2610.07023v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07023 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jinghao Pang [view email]
[v1]
Sun, 4 Oct 2026 15:50:42 UTC (1,991 KB)
来源:arXiv:cs.AI · arxiv.org