arXiv:cs.LG· Rodrigo Schuller, Francisco Ganacim·· 4 小时前AI 评分37
从第一性原理出发的数据集剪枝:一种无标签线性规划方法
Dataset Pruning from First Principles: A Label-Free Linear Programming Approach
AI 导读
研究者提出一种无标签、无需模型训练的数据集剪枝方法,将无偏子集选择重构为方差最小化问题,并把无偏子集选择算法刻画为高维多面体,用顶点游走近似优化方差目标。在 CIFAR-10、MNIST 和 CelebA 上,该方法在每个评测预算下的平均测试准确率均不逊于均匀采样,在小预算下优于同类几何方法,还能在批量不变的情况下通过提升 mini-batch 多样性降低随机梯度方差。
正文
Abstract:Dataset pruning reduces a large training set to a representative subset while preserving model performance. Existing geometry-based methods typically assume that nearby points in embedding space share similar properties. Rather than imposing this assumption, we derive geometric selection criteria by reformulating unbiased subset selection as a variance minimization problem. Unbiasedness ensures that unweighted subset averages recover full-dataset averages in expectation, including losses and gradients at fixed model parameters. Specifically, we characterize a family of unbiased subset selection algorithms as a high-dimensional polytope. In this context, minimizing the expected sampling variance is a linear objective. Differences in sampling variance, averaged over rigid motions, admit closed-form pairwise expressions. Because the polytope has high dimension, directly applying standard linear programming is impractical. We instead use these expressions to construct an efficient vertex walk that optimizes an approximation of the variance objective while preserving unbiasedness, yielding a method that requires neither labels nor model training during selection. Across CIFAR-10, MNIST, and CelebA benchmarks, our method matches or exceeds uniform sampling in mean test accuracy at every evaluated budget and outperforms competing geometric methods in several settings, particularly at small selection budgets. Beyond dataset pruning, the same framework reduces stochastic-gradient variance by increasing diversity within mini-batches while keeping the batch size unchanged.
| Subjects: | Machine Learning (stat.ML); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10347 [stat.ML] |
| (or arXiv:2610.10347v1 [stat.ML] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10347 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rodrigo Schuller [view email]
[v1]
Wed, 7 Oct 2026 16:24:50 UTC (3,574 KB)
来源:arXiv:cs.LG · arxiv.org