arXiv:cs.LG(机器学习,全量分类)· Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka·· 14 小时前AI 评分32
SAGE:基于相似度、利用少量已核验样本清洗被投毒训练数据
SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
AI 导读
针对数据投毒防御中大量已核验干净样本成本过高的问题,研究者提出 SAGE,仅依赖少量经专家核验的干净与投毒样本,在独立数据集上训练通用特征提取器,再用非参数、相似度加权的预测标记投毒样本。在针对七种 clean-label 攻击方法的标准基准上,即使只有少量已核验投毒样本也能带来显著优势,且已核验干净样本在各类别间的分布比其数量更重要。
正文
Abstract:As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.01788 [cs.LG] |
| (or arXiv:2610.01788v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01788 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Aron Laszka [view email]
[v1]
Thu, 1 Oct 2026 14:34:54 UTC (115 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org