arXiv:cs.AI· Zhankai Ye, Yanning Wang, Yukai Jin, Bo Mei, Fangyi Li, Wei Wang, Shangqian Gao, Xin Liu·· 3 小时前
叙事包装如何影响 LLM 拒答:跨语言基准与防御方法
How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
AI 导读
研究显示安全对齐模型在角色扮演或叙事包装下容易回答原本会拒绝的有害请求,Qwen3-1.7B 上英文攻击成功率达 89.4%,现代中文为 93.0%,文言文更达 95.7%。
正文
Abstract:Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11005 [cs.AI] |
| (or arXiv:2610.11005v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11005 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhankai Ye [view email]
[v1]
Wed, 7 Oct 2026 23:40:25 UTC (381 KB)
来源:arXiv:cs.AI · arxiv.org