arXiv:cs.LG(机器学习,全量分类)· Aayush Karan, Sitan Chen, Yilun Du·· 14 小时前AI 评分44
用采样做微调:SFT 的学习能力比你想的更强
Finetuning with Sampling: SFT Learns Better Than You Think
AI 导读
研究者提出一种 MCMC 采样算法,在给定参考模型下将 off-policy 轨迹逐步转化为更接近 on-policy 的数据,从而让 SFT 媲美主流后训练技术。在科学技能习得、数学推理和开放式专业知识等任务上,该方法常比强 on-policy 基线泛化更好、遗忘更少,并展现出超越锐化基座模型分布的分布性能。
正文
Abstract:Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02140 [cs.LG] |
| (or arXiv:2610.02140v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02140 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Aayush Karan [view email]
[v1]
Thu, 1 Oct 2026 17:45:07 UTC (158 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org