arXiv:cs.LG· Jaehui Hwang, Dongyoon Han, Sangdoo Yun, Byeongho Heo·· 4 小时前AI 评分33
负策略采样赋能在线策略蒸馏:NP-OPD 方法提出
On-Policy Distillation with Negative-Policy Rollouts
AI 导读
研究者提出 Negative-Policy OPD(NP-OPD),在在线策略蒸馏(OPD)的采样阶段引入能力更弱、表现更差的负策略,持续提供负策略偏好但教师不偏好的 token,让学生在接受教师正向监督的同时获得显式负向信号。
正文
Abstract:On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insufficient learning signals. In this work, we introduce Negative-Policy OPD (NP-OPD), which complements teacher supervision with rollouts from a lower-performing, lower-capability negative policy that serves as a negative reference for the student. Rather than modifying the distillation reward formulation, NP-OPD introduces the negative policy at the rollout stage, continuously supplying tokens preferred by the negative policy over the teacher so that they remain exposed to teacher supervision throughout training. This provides an explicit negative signal through negative-policy rollouts while preserving the positive teacher supervision used in OPD. Through extensive experiments, we show that NP-OPD improves OPD across model scales, generation modes, reasoning domains, and different OPD variants. Furthermore, our analyses show that NP-OPD effectively suppresses tokens preferred by the negative policy over the teacher and moves the student away from the negative policy. These results support our design of introducing negative signals through negative-policy rollouts and provide new insight into the role of the rollout policy in OPD. Code will be available at this https URL.
| Comments: | 25 pages, 7 figures, 24 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07874 [cs.LG] |
| (or arXiv:2610.07874v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07874 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jaehui Hwang [view email]
[v1]
Tue, 6 Oct 2026 07:25:03 UTC (357 KB)
来源:arXiv:cs.LG · arxiv.org