arXiv:cs.LG(机器学习,全量分类)· Yang Ba, Michelle V. Mancenido, Rong Pan·· 14 小时前AI 评分33
面向 LLM 微调的 Training-Aware Target Coverage 合成数据选择方法
Training-Aware Target Coverage for Synthetic Data Selection
AI 导读
研究者提出 Training-Aware Target Coverage(TATC)合成数据选择方法,用于 LLM 微调,通过线性理论刻画合成数据收益与误差的权衡,判断数据是否有用、该加多少及单条样本的边际价值。在数学推理任务中,TATC 为 Qwen2.5-Math-1.5B-Instruct 筛选合成解,在 GSM8K 上于各选择预算下均优于其他合成数据选择方法。
正文
Abstract:Synthetic data are increasingly used to scale LLM training, yet more synthetic data do not necessarily produce better models. Useful synthetic data must add information relevant to the target task without introducing errors that offset their benefit, and the value of an example can change as the training set grows. We develop a linear theory that characterizes this tradeoff and determines where synthetic data are useful, how much should be added, and the marginal value of adding one example to an existing set. The analysis shows the conditions when input coverage alone is sufficient and when synthetic errors must also be considered. Guided by these results, we introduce \emph{Training-Aware Target Coverage} (TATC), a synthetic data selection method for LLM fine-tuning. TATC identifies candidates whose training effects are beneficial to the target task and selects among them to expand coverage of target-relevant directions not already represented by the available data. Experiments on text and image data verify the linear theory. With mathematical reasoning tasks, TATC selects synthetic solutions for fine-tuning Qwen2.5-Math-1.5B-Instruct and outperforms alternative synthetic-data selection methods on GSM8K across selection budgets. In summary, we provide a principled approach to synthetic data selection by quantifying and maximizing its value to the target task.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00814 [cs.LG] |
| (or arXiv:2610.00814v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00814 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yang Ba [view email]
[v1]
Wed, 30 Sep 2026 23:17:32 UTC (191 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org