arXiv:cs.LG· Rui Li, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, Qi Liu·· 3 小时前AI 评分37
AIR-OPD:面向在线策略蒸馏的自适应迭代修复框架
Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
AI 导读
AIR-OPD 是一种面向在线策略蒸馏的自适应迭代修复框架,通过引导生成器针对学生模型当前错误合成修复引导,让学生进行在线策略重试,失败则针对新错误生成新引导。在 DAPO-Math-17K 上训练后,Qwen3-4B 和 Qwen3-8B 在 AIME24、AIME25、HMMT25 上的数学推理平均成绩较最强基线最高提升 3.6 分,同时在 MMLU-Pro 和 GPQA 上保持基座模型表现。
正文
Abstract:On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
| Comments: | 21 pages, 3 figures |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02700 [cs.LG] |
| (or arXiv:2610.02700v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02700 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rui Li [view email]
[v1]
Fri, 2 Oct 2026 02:34:57 UTC (197 KB)
来源:arXiv:cs.LG · arxiv.org