跳到正文
arXiv:cs.LG· Yifei Liu, Minghao Fang, Xinyu Gu, Chengkai Yao, Mengdi Liu, Tengfei Ma, Jiangbin Zheng, Chang Yu, Zhangyang Gao·· 4 小时前AI 评分35

E²-OPSD:解决在线策略自蒸馏中的熵超调问题

E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation

AI 导读

研究提出 E²-OPSD,通过示例引导教学和熵感知蒸馏两种机制,解决在线策略自蒸馏(OPSD)中出现的熵超调问题。在数学推理任务上,E²-OPSD 相比 OPSD 在 mean@16 上最高提升 4.3 分;域外评测中,相比对应基座模型在 mean@16 上最高提升 4.9 分、pass@8 上提升 5.5 分。该方法无需额外前向传播或网络,保持简洁。

正文

View PDF HTML (experimental)

Abstract:On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E$^2$-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E$^2$-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E$^2$-OPSD remains simple, requiring no additional forward passes or networks.
Comments: 24 pages, 6 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.05048 [cs.LG]
  (or arXiv:2610.05048v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.05048

arXiv-issued DOI via DataCite

Submission history

From: Yifel Liu [view email]
[v1] Sun, 4 Oct 2026 08:32:33 UTC (5,937 KB)
[v2] Tue, 6 Oct 2026 02:10:14 UTC (5,938 KB)

来源:arXiv:cs.LG · arxiv.org