arXiv:cs.LG· Kevin Kuo, Virginia Smith, Chhavi Yadav·· 4 小时前AI 评分58
arXiv 论文:开源权重 LLM 微调防护易受简单攻击,提出 ART 缓解方法
Open-Weight LLM Fine-Tuning Defenses are Susceptible to Simple Attacks
AI 导读
arXiv 论文(arXiv:2605.26526)指出,开源权重 LLM 的防护可被 abliteration 和 prefilling 两种低成本攻击绕过,在 BeaverTails、HarmBench 和 AdvBench 三个基准上,攻击成功率从低于 10% 提升到 16%-96%。
正文
Abstract:Recent defenses for safeguarding open-weight large language models (LLMs) are intended to prevent adversarial usage. Underlying these defenses is an assumption that new harmful behavior is learned through fine-tuning rather than elicited by jailbreaking the model. Yet, pretrained LLMs already encode substantial harmful knowledge across many domains, which raises an important question: can an adversary jailbreak safeguarded models, to achieve harmful usage without fine-tuning at all? In this paper, we show that open-weight safeguards are susceptible to simpler strategies that, despite being well known, have not been systematically evaluated against these safeguards. Specifically, we evaluate two low-cost attacks--abliteration and prefilling--that do not rely on gradient-based optimization. Across three harmfulness evaluation benchmarks (BeaverTails, HarmBench, and AdvBench), these attacks increase attack success rates against safeguarded open-weight models from below 10\% to a range of 16%-96%. To mitigate this vulnerability, we introduce abliteration-resistant tuning (ART), which incorporates an abliteration-based objective into training. ART can be layered onto existing defenses and reduces the success rates of abliteration, prefilling, and their combination by 10%-20%. These findings indicate that the attack surface for open-weight models is broader than previously characterized, and that evaluations of safeguarding defenses should incorporate a more diverse set of attack strategies beyond adversarial fine-tuning.
| Comments: | main body: 10 pages, 4 figures, 10 tables |
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2605.26526 [cs.LG] |
| (or arXiv:2605.26526v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.26526 arXiv-issued DOI via DataCite |
Submission history
From: Kevin Kuo [view email]
[v1]
Tue, 26 May 2026 04:18:42 UTC (874 KB)
[v2]
Tue, 6 Oct 2026 23:00:15 UTC (893 KB)
来源:arXiv:cs.LG · arxiv.org