arXiv:cs.CL· Chenglei Shen, Haoyang Yao, Weijie Yu, Song Jin, Xiao Zhang, Jun Xu·· 4 小时前AI 评分35
从修复后的推理中学习:根因引导的在线策略蒸馏
Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation
AI 导读
针对在线策略自蒸馏(OPSD)中参考解法无法解释学生自身推理错误、易致推理错配与蒸馏陷阱的问题,研究者提出根因引导的在线策略蒸馏(RC-OPD)。该方法对每次失败尝试定位最早实质性错误、生成局部修正,并以修正后的中间结果作为有效前缀锚点,通过迭代的诊断—修复—续写流程在固定修复预算内检验修复效果。多数据集与多模型规模的实验显示,RC-OPD 缓解了推理错配与蒸馏陷阱,取得显著性能提升。
正文
Abstract:On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student's own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis--repair--continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root--cause--guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03515 [cs.CL] |
| (or arXiv:2610.03515v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03515 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Chenglei Shen [view email]
[v1]
Fri, 2 Oct 2026 16:07:39 UTC (1,048 KB)
来源:arXiv:cs.CL · arxiv.org