arXiv:cs.LG· Lyuxin David Zhang, Eric Wong, Surbhi Goel, Anton Xue·· 3 小时前AI 评分36
LESSER:用输出层梯度做后训练数据选择
LESSER: Post-Training Data Selection with Output-Layer Gradients
AI 导读
LESSER 是一种基于输出层梯度的后训练数据选择方法,只需前向传播即可近似全参数梯度特征,将特征提取 FLOP 成本在 SFT 上降低 9.7 倍、在 RL 基准上降低 3.0 倍,同时保持下游任务性能与全梯度方法相当。研究发现,即使输出层梯度与全梯度对单样本的排序不同,二者选出的批次梯度仍是对齐的。
正文
Abstract:The choice of post-training data for large language models substantially affects downstream performance. Gradient-based data selection is a popular approach that ranks training data by how well their gradients align with those of a small validation set. However, ranking with full-parameter gradients requires an expensive backward pass on every sample, making computation intractable for large candidate pools. This raises a natural question: can we approximate full-gradient features at a fraction of the cost? Conveniently, we find that output-layer gradients suffice for effective data selection, yet require only the cheaper forward pass. We implement this as LESSER, a drop-in wrapper for selection methods that reduces the feature-extraction FLOP cost by $9.7\times$ for SFT and $3.0\times$ for RL benchmarks, while tracking full-gradient performance on downstream tasks. Empirically, we find that even when output-layer and full gradients rank individual samples differently, they select batches with aligned gradients.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03702 [cs.LG] |
| (or arXiv:2610.03702v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03702 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lyuxin David Zhang [view email]
[v1]
Fri, 2 Oct 2026 17:55:42 UTC (265 KB)
来源:arXiv:cs.LG · arxiv.org