跳到正文
arXiv:cs.AI· Yafeng Tang, Hao Li, Hongsheng Yu, Qiang Fu·· 3 小时前

ReTeach:通过多轮反思与重试构建自教师模型

ReTeach: Building a Self-Teacher through Multi-Round Reflection and Retry

AI 导读

ReTeach 是一个反思式自蒸馏框架,仅凭模型自身生成的尝试和结果级验证,通过多轮反思与重试构建自教师,无需参考答案或外部反馈。在数学推理、科学问答和工具使用等六个基准上,其平均准确率较 GRPO 提升 1.39 个百分点。该方法采用结果感知的选择与加权策略,区分初始正确、反思纠正和未解决样本,并通过 on-policy 蒸馏让学生模型在单次推理下获得迭代纠正的收益。

正文

View PDF HTML (experimental)

Abstract:Self-distillation can improve reasoning without a separately trained, more capable teacher, but its effectiveness depends on how the self-teacher gains an advantage over the student. Conditioning the teacher on reference answers or solutions can provide such an advantage, but this information may be unavailable. Reflection offers a way to derive explicit error diagnoses and revision guidance from self-generated attempts, yet existing reflection-based methods often combine it with reference information, rich task feedback, or persistent memory. We introduce ReTeach, a Reflective self-distillation framework that constructs its self-Teacher through multi-round reflection and retry using only self-generated attempts and outcome-level verification. Starting from an unsuccessful student rollout, the teacher alternates explicit reflection with renewed attempts until success or the retry budget is exhausted, without reference answers or solutions, external diagnostic feedback, or cross-example memory. Each failed retry informs subsequent reflection, while successful correction provides outcome-level evidence for the potential utility of the resulting teacher context. An outcome-aware selection and weighting strategy distinguishes initially correct, reflection-corrected, and unresolved examples, assigning separate weights to their category-normalized distillation losses. Through on-policy distillation, the student matches the teacher's context-conditioned token-level predictive distributions at prefixes of its own rollouts, transferring the benefits of iterative correction while retaining single-pass inference. Across six benchmarks spanning mathematical reasoning, science question answering, and tool use, ReTeach improves average accuracy over GRPO by 1.39 percentage points.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11529 [cs.AI]
  (or arXiv:2610.11529v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11529

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yafeng Tang Mr. [view email]
[v1] Thu, 8 Oct 2026 08:58:01 UTC (380 KB)

来源:arXiv:cs.AI · arxiv.org