无数据在线策略蒸馏(DF-OPD):不依赖外部数据能走多远?
Data-Free On-Policy Distillation: How Far Can We Go Without External Data?
研究发现在线策略蒸馏(OPD)对训练数据的依赖远低于预期:仅用 8 条真实提示词的训练效果即可媲美 17k 道题。作者据此提出无数据在线策略蒸馏(DF-OPD),无需种子示例自生成 64 个问题即可达到全量数据 OPD 的水平;在多教师场景下,1k 个生成问题可媲美约 7k 条真实后训练样本。
Authors:Gengsheng Li, Mao Zheng, Mingyang Song, Jie Sun, Zeyuan Liu, Ruiqi Liu, Tianyu Yang, Qiyong Zhong, Haiyun Guo, Junfeng Fang, Shiming Xiang, Jinqiao Wang, Tat-Seng Chua
Abstract:On-policy distillation (OPD) is increasingly applied to frontier foundation model post-training. Prior work in this area has largely focused on algorithmic advances, yet it remains unclear how much OPD depends on its training questions and, in particular, how far this dependence can be reduced. Across two representative single-teacher OPD settings, we find that training on 8 real prompts yields performance comparable to training on 17k problems, while datasets differing substantially in measured difficulty and initial distillation gap yield similar outcomes. Our analyses suggest two complementary explanations: repeated sampling could allow even a few prompts to expose substantial teacher supervision, while OPD transfers generalizable reasoning capabilities beyond dataset-specific knowledge. Building on these observations, we next propose a data-free on-policy distillation (DF-OPD) setting to investigate whether the system can supply the training questions itself, eliminating the need for external data. With 64 self-generated questions obtained without seed examples, DF-OPD yields performance comparable to full-data OPD in both single-teacher settings. This finding also holds in multi-teacher OPD: across mathematics, code, and instruction following, 1k generated questions achieve performance comparable to training on approximately 7k real post-training examples. We further explore whether OPD can operate even without explicit training questions. The experiments show that this is effective only in limited cases, where the student unexpectedly generates and answers its own questions, thereby reducing the process to an implicit form of DF-OPD. Together, these findings invite a reassessment of the role of training data in on-policy distillation. Code is available at this https URL
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.14193 [cs.LG] |
| (or arXiv:2609.14193v4 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.14193 arXiv-issued DOI via DataCite |
Submission history
From: Gengsheng Li [view email]
[v1]
Sat, 12 Sep 2026 23:51:35 UTC (320 KB)
[v2]
Thu, 17 Sep 2026 20:45:40 UTC (453 KB)
[v3]
Thu, 1 Oct 2026 20:06:23 UTC (1,155 KB)
[v4]
Wed, 7 Oct 2026 17:57:03 UTC (1,160 KB)
来源:arXiv:cs.AI · arxiv.org