arXiv:cs.AI(全量分类)· Jianwei Li, Min-Seon Kim, Jung-Eun Kim·· 5 小时前AI 评分44
QES:通过专家隔离与关停实现 LLM 后门遏制
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
AI 导读
研究者提出 Quarantined Expert Shutdown(QES),一种在类 MoE 结构中通过正则化引导的后门遏制策略:训练时允许后门形成,但将其路由至可隔离的指定专家,部署时只需将隔离专家的路由权重置零即可完成缓解。在两项任务、三种攻击和四个模型家族上,该方法将攻击成功率从 100% 降至 0-10%,下游效用通常保持或仅受轻微影响。
正文
Abstract:Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a computationally efficient containment strategy built in a regularization-steered MoE-like setting. Specifically, given a poisoned dataset, QES augments a Transformer-based language model with routed expert-specific LoRA branches and lightweight routers, and uses auxiliary routing objectives to attract trigger-conditioned behavior into a designated expert while preserving benign capability elsewhere. At deployment, mitigation reduces to a single constant-time operation: zeroing the quarantined expert's routing weight, without trigger screening or further updating model weights. Empirically, our methods reduce the attack success rate ASR from 100% to 0-10% on most settings across two tasks, three attacks, and four model families, while downstream utility is often preserved or only modestly affected. These results establish learn, but channel as a previously unexplored regime for backdoor containment in generative LLMs.
| Comments: | NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00663 [cs.AI] |
| (or arXiv:2610.00663v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00663 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jung-Eun Kim [view email]
[v1]
Wed, 30 Sep 2026 20:02:14 UTC (61 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org