跳到正文
原文
arXiv:cs.AI(全量分类)· Jianwei Li, Jung-Eun Kim·· 5 小时前AI 评分51

基于零空间投影的 LoRA 微调大语言模型后门净化方法

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection

AI 导读

arXiv 论文(NeurIPS 2026,arXiv:2610.00685)提出一种针对 LoRA 微调 LLM 的后门净化方法,无需触发器先验、干净参照或事后重训练。方法通过数据筛选与特征近似提取后门方向,在每层或每个头的输入和输出通道构建正交零空间,将 LoRA 更新投影其上。实验显示攻击成功率(ASR)从接近 100% 降至 10% 以下,同时保留基座模型的通用能力和适配器学到的下游技能。

正文

View PDF HTML (experimental)

Abstract:With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general capabilities and (ii) the new downstream skills learned through the adapter. Through a series of ablation studies, we progressively scale our approach from a single layer in a text classification setting to a full-parameter LLM in the generative task. Through careful data curation and feature approximation, we extract high-fidelity backdoor directions and, for each layer or head, construct orthogonal null spaces in both the input and output channels, onto which the LoRA updates are projected. Empirically, our null-space projection method reduces the ASR from nearly 100% to less than 10%, while preserving the base model's benign performance and the adapter's learned abilities during downstream task adaptation.
Comments: NeurIPS 2026
Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as: arXiv:2610.00685 [cs.AI]
  (or arXiv:2610.00685v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00685

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jung-Eun Kim [view email]
[v1] Wed, 30 Sep 2026 20:28:34 UTC (364 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org