跳到正文
arXiv:cs.LG· Itay Zloczower, Eyal Lenga, Gilad Gressel, Yisroel Mirsky·· 5 小时前AI 评分66

研究显示针对恶意微调的对齐防御在持续训练下会失效

A Few Steps Further: Why Defenses Against Malicious Finetuning Erode Under Continued Training

AI 导读

arXiv 论文测试了六种防御在四个开源权重模型上的表现,将同一有害数据微调攻击延长到三个 epoch 后,全部 72 次防御运行结束时模型有害性都高于发布时,Llama-3.1 在最高学习率下防御模型几乎与无防御模型一样有害。作者指出十五种近期防御共享同一弱点,即都建立在有限的攻击者模型上,权重发布后保护无法持续,目前不应视为可靠防护。

正文

View PDF HTML (experimental)

Abstract:Model providers increasingly release the weights of large language models. Although these models are safety-aligned before release, their safeguards can often be removed by fine-tuning on harmful data. A growing class of defenses aims to make alignment robust to such malicious fine-tuning, but these defenses are typically evaluated against attacks with a fixed training budget, even though an attacker who holds the weights can simply train for longer. We ask whether current defenses withstand this simplest escalation. Surveying fifteen recent defenses, we find that they share a common weakness: each is built around a limited model of the attacker, such as a bounded perturbation, a short simulated attack, or a trained link between harmful and benign behavior, and nothing enforces that protection once the weights are released. We then test six representative defenses on four open-weight models by continuing the same harmful-only fine-tuning attack for three epochs and measuring harmfulness and capability along the way. In all 72 defended runs, the model is more harmful at the end of training than at release, and on Llama-3.1 at the highest learning rate the defended models end almost as harmful as the undefended one. The defenses are not equally weak: one defense kept harmfulness low on one model, and some attacks recovered harmfulness only at the cost of general capability. Current defenses can delay or disrupt malicious fine-tuning, but in most cases their measured resistance does not persist under continued training, and they should not yet be treated as durable protection.
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2605.14605 [cs.CR]
  (or arXiv:2605.14605v3 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2605.14605

arXiv-issued DOI via DataCite

Submission history

From: Itay Zloczower [view email]
[v1] Thu, 14 May 2026 09:22:14 UTC (199 KB)
[v2] Sun, 24 May 2026 08:34:13 UTC (199 KB)
[v3] Wed, 7 Oct 2026 16:55:56 UTC (265 KB)

来源:arXiv:cs.LG · arxiv.org