跳到正文
arXiv:cs.AI· Yuxiang Zhang, Ding Cao, Shuting Cui, Lei Wang, Weijieying Ren, Tianxiang Zhao·· 5 小时前AI 评分35

Seg-OPD:用分段在线策略蒸馏让大语言模型学会修正推理

Learning to Revise Reasoning with Segment-wise On-Policy Distillation

AI 导读

研究者提出分段在线策略蒸馏(Seg-OPD),通过不确定性指标选出学生模型的推理片段并获取教师重写版本,训练学生偏好教师重写而非自身片段,同时保留密集的 token 级在线蒸馏监督。在数学推理与竞赛编程任务上,Seg-OPD 训练的模型修正成功率高于基线,推理准确率平均相对提升 5.22%。代码已开源。

正文

View PDF HTML (experimental)

Abstract:On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revised to improve subsequent reasoning. Furthermore, this paradigm can become less effective when the student produces a degenerate reasoning prefix, as subsequent teacher supervision remains conditioned on that prefix and may reinforce poor reasoning patterns. In this work, we focus on learning reasoning revision with segment-wise OPD to rework intermediate reasoning steps and better support subsequent reasoning. Through controlled reasoning interventions, we find that replacing student segments with teacher redrafts improves subsequent reasoning accuracy. Therefore, we address the problem of turning teacher redrafts into explicit supervision for learning to revise reasoning. We propose Segment-wise On-Policy Distillation (Seg-OPD), which selects student segments based on an uncertainty metric and obtains corresponding teacher redrafts. Seg-OPD trains the student to prefer teacher redrafts over their paired student segments while retaining dense token-wise OPD supervision. Extensive experiments on mathematical reasoning and competitive programming tasks show that Seg-OPD-trained students achieve higher revision success rates than baselines. Seg-OPD consistently outperforms the compared state-of-the-art baselines in reasoning accuracy with an average relative improvement of 5.22% across diverse models and tasks. Code is available at this https URL.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02703 [cs.AI]
  (or arXiv:2610.02703v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02703

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuxiang Zhang [view email]
[v1] Fri, 2 Oct 2026 02:40:14 UTC (6,517 KB)

来源:arXiv:cs.AI · arxiv.org