arXiv:cs.CL· Siyan Zhao, Yonggan Fu, Jindong Jiang, Shih-Yang Liu, Song Bian, Byung-Kwan Lee, Sharath Turuvekere Sreenivas, Wenliang Dai, Hanrong Ye, Aditya Grover, Pavlo Molchanov·· 3 小时前
何时需要 On-Policy Distillation?基于离线学生 rollout 的蒸馏往往更好
When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
AI 导读
研究提出 Semi-OPD,用初始学生生成的离线 rollout 进行蒸馏,在 1.5B 至 235B 参数的 17 组师生配对中有 14 组优于 On-Policy Distillation(OPD),准确率最高提升 13.6%,训练提速 11.4 倍。
正文
Authors:Siyan Zhao, Yonggan Fu, Jindong Jiang, Shih-Yang Liu, Song Bian, Byung-Kwan Lee, Sharath Turuvekere Sreenivas, Wenliang Dai, Hanrong Ye, Aditya Grover, Pavlo Molchanov
Abstract:On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11291 [cs.CL] |
| (or arXiv:2610.11291v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11291 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Siyan Zhao [view email]
[v1]
Thu, 8 Oct 2026 05:54:03 UTC (1,047 KB)
来源:arXiv:cs.CL · arxiv.org