跳到正文
arXiv:cs.AI· Andrew Draganov, Tolga H. Dur, Anandmayi Bhongade, Mary Phuong·· 3 小时前

Phantom Transfer:数据投毒可绕过数据级防御

Phantom Transfer: Data Poisoning can Survive Data-Level Defences

AI 导读

研究者提出名为 Phantom Transfer 的数据投毒攻击,即使已知毒样本如何被植入良性数据集也无法将其过滤,且攻击效果与数据生成模型、被训练模型及攻击目标无关。该攻击可绕过 11 种已测试的数据级防御,包括用另一模型对每个样本进行改写,并能向模型植入密码触发的行为。论文建议未来防御应补充白盒方法与训练后模型审计,该工作已被 NeurIPS 2026 接收。

正文

View PDF

Abstract:We present a data poisoning attack -- Phantom Transfer -- with the property that, even if you know precisely how the poison was placed into an otherwise benign dataset, you cannot filter it out. We achieve this by modifying subliminal learning to work in real-world contexts and demonstrate that the attack works regardless of which model produced the data, which model is trained on the data or what the attack target is. Furthermore, the attack survives 11 tested data-level defences, including one where every sample is paraphrased by another model. We characterise when this attack works best and show that it can be used to plant password-triggered behaviours into models while still beating defences. We suggest that future defences should be supplemented with white-box methods and post-training model audits.
Comments: Camera-ready version accepted at NeurIPS 2026; expanded experiments and model audits
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.04899 [cs.CR]
  (or arXiv:2602.04899v3 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2602.04899

arXiv-issued DOI via DataCite

Submission history

From: Tolga Dur [view email]
[v1] Tue, 3 Feb 2026 14:38:07 UTC (325 KB)
[v2] Tue, 2 Jun 2026 15:48:30 UTC (498 KB)
[v3] Wed, 7 Oct 2026 21:35:49 UTC (512 KB)

来源:arXiv:cs.AI · arxiv.org