跳到正文
arXiv:cs.LG· Martin V. Vejling, Shashi Raj Pandey, Christophe A. N. Biscio, Petar Popovski·· 3 小时前

面向分布内数据获取的保形数据污染测试

Conformal Data Contamination Tests for In-distribution Data Acquisition

AI 导读

研究者提出一种免分布假设、感知数据污染的数据获取框架,仅需检查少量样本即可识别哪些外部数据智能体的数据对模型个性化最有价值。该框架引入保形数据污染测试,在任意污染水平下保持有效性,其中 Storey 型测试可通过 Benjamini-Hochberg 过程实现有限样本下的错误发现率控制。在多种协作学习场景中的实验验证了该方法的稳健性与有效性。

正文

View PDF

Abstract:The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. The set of quality data can be expanded through trading or sharing with external data agents. However, external data may be contaminated or introduce undesirable sample diversity which can degrade performance of personalized machine learning tasks, as in diagnosis of a rare disease or recommendation systems. Therefore, data buyers need quality guarantees prior to data acquisition. Previous works primarily rely on distributional assumptions about data from different agents, relegating quality checks to post-hoc steps involving costly data valuation procedures. We propose a distribution-free, contamination-aware data acquisition framework that, by inspecting only a small volume of data, identifies external data agents whose data is most valuable for model personalization. To achieve this, we introduce novel two-sample testing procedures, preceding full data acquisition, grounded in rigorous theoretical foundations for conformal outlier detection, to determine whether an agent's data exceeds a contamination threshold. The proposed tests, termed conformal data contamination tests, remain valid under arbitrary contamination levels and the novel Storey-type test provably enables finite-sample false discovery rate control via the Benjamini-Hochberg procedure. Empirical evaluations across diverse collaborative learning scenarios demonstrate the robustness and effectiveness of our approach. Overall, the conformal data contamination test distinguishes itself as a generic procedure for aggregating data with statistically rigorous quality guarantees.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG)
Cite as: arXiv:2507.13835 [stat.ML]
  (or arXiv:2507.13835v2 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2507.13835

arXiv-issued DOI via DataCite

Submission history

From: Martin Voigt Vejling [view email]
[v1] Fri, 18 Jul 2025 11:44:42 UTC (452 KB)
[v2] Thu, 8 Oct 2026 09:06:24 UTC (528 KB)

来源:arXiv:cs.LG · arxiv.org