arXiv:cs.AI· Ruizhi Xu, Wei Xu, Sibo Zhu·· 6 小时前AI 评分48
TARE:用未中毒孪生模型重新度量后门防御的真实代价
TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
AI 导读
研究提出 TARE,用同一配方、调度和种子训练的未中毒孪生模型来度量后门防御的"剥离代价"(tare);发现 BackdoorBench 上 WaNet、BPP、Input-Aware 的攻击配置中 MultiStepLR 从未触发,导致受害模型从未退火、在 30/31 个公开 CIFAR 单元中准确率最低(≤5%)。
正文
Abstract:Backdoor-defense leaderboards print a clean-accuracy drop and read it as removal cost. Measured on the poisoned victim alone, the drop cannot separate removal from what the defense does to any model, and inherits the victim's start, which for three of BackdoorBench's sixteen attacks is a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR that never fires, so their victims never anneal and are the least accurate in 30/31 public CIFAR cells at $\leq$5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to their own level, so there the published cost is negative, the benchmark's rating clips the "gain" to zero, and 2 of 48 citing defense papers we read rest a no-cost claim on those cells; TSBD and CGD, re-run with their code, "gain" on a never-poisoned model too. A $2\times2$ editing only that scheduler line isolates the cause, its swapped arms self-registered before they ran: the sign of the fine-tuning family's clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. TARE runs the same defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, $\leq$10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On the BadNets grid seven of eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 (fine-tuning) to +5.70 points (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share unidentified. We ship the three-key patch, a signed tare column (7 attacks $\times$ 8 defenses) and TARE-Z, a twin-free estimator for seed-stable defenses.
| Comments: | 84 pages (9-page main text) |
| Subjects: | Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06994 [cs.CR] |
| (or arXiv:2610.06994v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06994 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sibo Zhu [view email]
[v1]
Sun, 4 Oct 2026 07:42:42 UTC (459 KB)
来源:arXiv:cs.AI · arxiv.org