跳到正文
arXiv:cs.LG· Jonas J\"ur{\ss}, Pietro Li\`o·· 5 小时前AI 评分35

如何识别导致 Phantom Transfer 的数据集偏差

Towards Identifying the Dataset Biases Causing Phantom Transfer

AI 导读

研究者提出一种识别数据集偏差的方法:用 Sentence-BERT 嵌入数据集补全结果,减去干净参考补全的嵌入向量,再与开放候选主题词表比对。在已知攻击者教师模型时,该方法识别偏差主题的 Matthews 相关系数达 0.83,未知时为 0.46。研究还发现不同教师模型会通过不同词汇表达同一偏差。

正文

View PDF HTML (experimental)

Abstract:Recent work has shown that a teacher model can transfer a bias to a student through a dataset from which every explicit reference to that bias has been filtered out, and that none of the tested data-level defenses reliably removes or detects such a bias, even when the defender knows what to look for. To shed light on the hidden traces these biases leave, we embed a dataset's completions with Sentence-BERT, subtract the embeddings of clean reference completions, and compare the result to an open vocabulary of candidate topics. This simple signature identifies the topic of the bias with a Matthews correlation coefficient of 0.83 when the attacker's teacher model is known, and 0.46 when it is not. We also observe that different teacher models appear to express the same bias through different vocabulary.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.14449 [cs.LG]
  (or arXiv:2609.14449v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.14449

arXiv-issued DOI via DataCite

Submission history

From: Jonas Jürß [view email]
[v1] Sun, 13 Sep 2026 11:43:20 UTC (52 KB)
[v2] Fri, 2 Oct 2026 01:23:26 UTC (59 KB)

来源:arXiv:cs.LG · arxiv.org