跳到正文
arXiv:cs.LG· Jian Luo, Kehan Qi, Qingqiao Hu, Meilong Xu, Jiacheng Qiu, Weimin Lyu, Jiawei Zhou, Chao Chen·· 4 小时前AI 评分39

安全对齐的 on-policy 蒸馏会传播后门吗?3% 投毒率可达 70% 攻击成功率

Does On-Policy Distillation for Safety Pose Backdoor Risks?

AI 导读

研究发现,安全对齐但被植入后门的教师模型可通过 on-policy 蒸馏(OPD)将隐藏恶意行为传给原本干净的学生模型,投毒率低至 3% 时学生攻击成功率(ASR)最高达 70%。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) has attracted growing attention as an effective way to transfer capabilities from teacher models to student models. Recent studies further explore OPD as a tool for improving large language model safety with promising results. However, these approaches typically assume that the teacher and training data are trustworthy. In this paper, we uncover an overlooked threat to OPD for safety: a safety-aligned but backdoored teacher can propagate its hidden malicious behavior to an initially clean student. Under our threat model, a poisoning rate as low as 3% results in an attack success rate (ASR) of up to 70% on the distilled student. We further identify two training choices that can amplify this risk. First, increasing the number of training epochs can lead to high ASR even at low poisoning rates. With only 10 poisoned samples, ASR reaches 67% after 16 epochs. Second, the commonly used top-k KL can accelerate backdoor transfer, causing trigger-conditioned harmful behavior to emerge earlier than sampled-token KL in most settings. Alongside these findings, we explore a simple mitigation, Lazy Defense, which clips KL rewards to make student updates less aggressive, limiting aggressive updates and slowing backdoor learning. Experiments show that Lazy Defense delays backdoor transfer in low poisoning rate settings. Together, our findings reveal that OPD can propagate backdoors, highlighting the need to address the safety risks of OPD.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07654 [cs.LG]
  (or arXiv:2610.07654v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07654

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jian Luo [view email]
[v1] Tue, 6 Oct 2026 02:50:07 UTC (373 KB)

来源:arXiv:cs.LG · arxiv.org