跳到正文
arXiv:cs.CL· Anhao Zhao, Haoran Xin, Junlong Tong, Yingqi Fan, Xuan Lu, Ping Nie, Wenjie Li, Xiaoyu Shen·· 4 小时前

DIAL-OPD:在 On-Policy 蒸馏中用更少 token 学到更多

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

AI 导读

DIAL-OPD 是一种 token 选择方法,通过用教师与学生概率的对数均值加权奖励幅度,在 on-policy 蒸馏中只保留 40% 的 token 即可超过 Vanilla OPD 及其全 token 变体,平均准确率最高提升 5.25 个百分点。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11659 [cs.CL]
  (or arXiv:2610.11659v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.11659

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Anhao Zhao [view email]
[v1] Thu, 8 Oct 2026 10:36:00 UTC (1,671 KB)

来源:arXiv:cs.CL · arxiv.org