arXiv:cs.AI· Bingxi Hou, Guochao Jiang, Guofeng Quan, Weiqing Li, Wenfeng Feng, Guohua Liu, Yuewei Zhang·· 6 小时前AI 评分36
重新审视跨 Tokenizer 的 On-Policy Distillation:从对齐覆盖到监督可靠性
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
AI 导读
研究挑战了"扩大对齐覆盖即可提升跨 tokenizer On-Policy Distillation 效果"的假设:在三组异构师生模型(数学推理与代码生成)上,严格 1:1 对齐位置已覆盖学生生成的大部分 token,将反向 KL 限制在共享词表中学生选定的 top-16 子集即可达到全共享词表 OPD 的准确率,并优于所评估的跨 tokenizer 基线。
正文
Abstract:On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08448 [cs.CL] |
| (or arXiv:2610.08448v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08448 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Guochao Jiang [view email]
[v1]
Tue, 6 Oct 2026 14:37:22 UTC (734 KB)
来源:arXiv:cs.AI · arxiv.org