跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yangze Liu, Xiao-Long Yin, Zhongyi Han·· 14 小时前AI 评分38

研究:样本价值只相对于学习器而定义,ResNet-18 加宽使 crossover 从 57 移至 85 样本/类

Useful to Whom? Sample Value Is Defined Only Relative to the Learner

AI 导读

一项受控实验表明,coreset 选择中 easy-first 与几何覆盖准则的 crossover 边界并非由数据固定,而是随目标学习器变化:在低分辨率 ImageNet-100 上,将 ResNet-18 宽度加倍使 crossover 从每类 57 个样本移至 85 个。

正文

View PDF HTML (experimental)

Abstract:What kind of data does a model need in order to learn? Coreset selection makes this question concrete: under a budget, keep the samples most useful for training. Easy-first and geometric coverage criteria can win in different budget regimes, separated by a crossover boundary. We ask whether this boundary is fixed by the data or changes with the target learner. Controlled experiments freeze the selected subsets and manipulate only the training learner. On low-resolution ImageNet-100, doubling ResNet-18's width moves the crossover from 57 to 85 samples per class: the learner changes the relative value of the same samples. A wider sweep reveals an interaction between input grid and capacity. Enlarging the grid while retaining the same image information shifts the boundary left, and this shift weakens as width increases. Stride controls reproduce and reverse the grid effect without changing the input grid; removing only the last downsampling stride is sufficient to recover the leftward shift. Under the native-224px ImageNet-1k protocol, width effects are smaller and depend on the probe: LFrac remains nearly flat, while EL2N shifts modestly right. Swapping the convolutional learning system for a ViT makes coverage win throughout the measured range, even when the easy subsets come from the convolutional proxy. These results establish learner dependence through frozen-subset interventions and identify network structure that can move the boundary. They do not yield a universal scaling law. Their practical implication is direct: a selection strategy's preferred budget regime must be evaluated with respect to the target learner.
Comments: 21 pages
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00221 [cs.LG]
  (or arXiv:2610.00221v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00221

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yangze Liu [view email]
[v1] Tue, 22 Sep 2026 05:17:34 UTC (185 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org