跳到正文
arXiv:cs.LG· Hongyu Chen, Xinyi Luo, Ming Zhao, Lin Tang, Zihan Xu, Jing Li, Yuxuan Wang, Haoran Deng, Wei Zhang·· 2 天前AI 评分38

LOOM:用留一法梯度匹配实现 LLM 微调的在线数据选择

Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning

AI 导读

针对在线批次选择中样本内梯度匹配反而偏向噪声样本、难以超越全批次训练的问题,研究者提出 LOOM——在 Adam 预条件子度量下于常规反向传播中计算 Gram 矩阵,去掉对角线后按梯度信噪比加权贪心选子集,无需留出数据。

正文

View PDF HTML (experimental)

Abstract:Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the data distribution, so removing the diagonal turns the matching objective into an unbiased estimate of the update's error with respect to the population gradient. The minimizer of this leave-one-out objective weights examples by their gradient signal-to-noise ratio (SNR), and whenever per-example SNR is heterogeneous enough, half of a batch yields a lower-error update than the whole batch; we give the exact condition. We build \method{} on this principle. It computes the Gram matrix in the metric of the Adam preconditioner during the ordinary backward pass, selects a weighted subset greedily with a $(1-e^{-\gamma})$ guarantee, and uses no held-out data. Across four fine-tuning tasks and seven backbones from 1.5B to 8B parameters, LOOM improves on full-batch training by 2.3 and 2.4 points on Llama-3.1-8B and Qwen2.5-7B, exceeds every in-sample gradient matcher by 2.4 points and the validation-guided GREATS and OPUS by 1.6--2.0, and selects injected label noise at under a fifth of its base rate.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00436 [cs.LG]
  (or arXiv:2610.00436v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00436

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Lin Tang [view email]
[v1] Wed, 30 Sep 2026 17:22:54 UTC (854 KB)

来源:arXiv:cs.LG · arxiv.org