跳到正文
arXiv:cs.LG· Jiacheng Liang, Yuhui Wang, Tanqiu Jiang, Ting Wang·· 4 小时前AI 评分39

RASA:面向 Mixture-of-Experts 模型的路由感知安全对齐框架

Routing-Aware Safety Alignment for Mixture-of-Experts Models

AI 导读

研究者提出路由感知的专家级对齐框架 RASA,针对 MoE 模型稀疏路由导致的安全对齐难题,通过识别被成功越狱不成比例激活的专家、在固定路由下仅微调这些专家并强制路由与安全对齐上下文保持一致。在两种代表性 MoE 架构和多种越狱攻击下,RASA 实现近乎完美的鲁棒性和跨攻击泛化能力,同时显著降低过度拒答,并在 MMLU、GSM8K、TruthfulQA 上保持通用能力。

正文

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) language models introduce unique challenges for safety alignment due to their sparse routing mechanisms, which can enable degenerate optimization behaviors under standard full-parameter fine-tuning. In our preliminary experiments, we observe that naively applying full-parameter safety fine-tuning to MoE models can reduce attack success rates through routing or expert dominance effects, rather than by directly repairing Safety-Critical Experts. To address this challenge, we propose RASA, a routing-aware expert-level alignment framework that explicitly repairs Safety-Critical Experts while preventing routing-based bypasses. RASA identifies experts disproportionately activated by successful jailbreaks, selectively fine-tunes only these experts under fixed routing, and subsequently enforces routing consistency with safety-aligned contexts. Across two representative MoE architectures and a diverse set of jailbreak attacks, RASA achieves near-perfect robustness, strong cross-attack generalization, and substantially reduced over-refusal, while preserving general capabilities on benchmarks such as MMLU, GSM8K, and TruthfulQA. Our results suggest that robust MoE safety alignment benefits from targeted expert repair rather than global parameter updates, offering a practical and architecture-preserving alternative to prior approaches.
Comments: 9 pages
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)
Cite as: arXiv:2602.04448 [cs.LG]
  (or arXiv:2602.04448v4 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2602.04448

arXiv-issued DOI via DataCite

Submission history

From: Jiacheng Liang [view email]
[v1] Wed, 4 Feb 2026 11:19:15 UTC (1,598 KB)
[v2] Sat, 4 Apr 2026 15:11:27 UTC (1,519 KB)
[v3] Mon, 5 Oct 2026 07:24:02 UTC (851 KB)
[v4] Wed, 7 Oct 2026 07:27:23 UTC (851 KB)

来源:arXiv:cs.LG · arxiv.org