arXiv:cs.CL· Xiaobing Chen, Zhiqi Pang·· 3 小时前
Residual Advantage:面向可验证奖励 RL 的学生相对教师引导方法
Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
AI 导读
研究者提出 Residual Advantage(RA),将教师与学生概率残差视为有界单步奖励,减去学生策略下的状态价值形成标准优势,并在每条回答内中心化后叠加到验证器优势上,使验证器优势仍保持为回答的均值标签。
正文
Abstract:Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step reward, subtracts the corresponding state value under the student policy to form a standard advantage, and centers the result within each response before adding it to the verifier advantage. The guidance term has zero mean within each response, so the verifier advantage remains the response's mean label and the teacher only redistributes credit among the steps within it. \CoRA{} further updates a teacher LoRA with verifier advantages on the same scored student batch and uses the updated teacher in the next iteration's residual, adapting guidance to the student's attempts. With Qwen3-1.7B-Base and Qwen3-4B-Base students and a Qwen3-8B teacher, \RA{} combined with GRPO or REINFORCE++ improves the underlying sequence-advantage algorithm in all 24 comparisons on three mathematical benchmarks, raising macro Avg@8 by 1.7--3.6 points and Pass@8 by 3.9--6.3 points. Both combinations surpass teacher-only OPD, and \CoRA{} adds a further 1.0--1.5 Avg@8 points.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11519 [cs.CL] |
| (or arXiv:2610.11519v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11519 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xiaobing Chen [view email]
[v1]
Thu, 8 Oct 2026 08:52:55 UTC (570 KB)
来源:arXiv:cs.CL · arxiv.org