arXiv:cs.LG(机器学习,全量分类)· Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim·· 14 小时前AI 评分47
拒绝行为可定位,破坏却会转移:少样本微调下的安全层研究
Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
AI 导读
研究测试了对 LLM 安全行为的定位与修复能否抵御攻击变化,覆盖四个模型家族六个 checkpoint。冻结至可复现的拒绝恢复深度后,100 个有害样本仍使全部六个 checkpoint 的拒绝率接近零;在 Llama-3.1-8B 上,常规训练会削弱修复,攻击者分散更新即可破解,基于良性微调校准的谱检测器也漏检多数修复失败。
正文
Abstract:Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization and recovery support defenses that survive changes in the attack. Across six checkpoints from four model families, harmful and benign prompts remain linearly separable after attack, and patching full clean hidden states into the compromised model restores refusal at a reproducible transition depth. Building on a prior layer-freezing defense, we freeze every layer up to this depth and repeat the attack. At a hundred harmful examples, refusal remains near zero on all six checkpoints, with recovery transitions above the frozen boundary. In a second study, removing the update's top two singular directions restores refusal after short attention-only fine-tunes on four checkpoints. On Llama-3.1-8B, ordinary training changes weaken this repair and an attacker who spreads the update defeats it. A spectral detector calibrated on benign Llama fine-tunes misses most repair failures on that checkpoint. Localized freezing can nevertheless help preserve refusal when a few harmful examples enter training data unintentionally. These results show that an attacker can bypass a region identified by recovery and defeat a repair that works across multiple checkpoints, motivating five checks for defenses against adaptive fine-tuning. Code is available at this https URL.
| Comments: | 24 pages, 7 figures, 21 tables. Jungseob Lee and Dongyub Jude Lee contributed equally |
| Subjects: | Computation and Language (cs.CL); Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00320 [cs.CL] |
| (or arXiv:2610.00320v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00320 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jungseob Lee [view email]
[v1]
Tue, 29 Sep 2026 04:12:28 UTC (308 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org