跳到正文
arXiv:cs.LG· Xinpeng Wang, Wei Shi, Yu-Chia Chen, Maria Zontak, Yun He, Richard Yuanzhe Pang·· 3 小时前AI 评分36

OPD Before RL:用同策略蒸馏为基于评分标准的 RL 预热

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

AI 导读

研究提出两阶段训练框架,先用评分标准作为教师上下文做同策略蒸馏(RP-OPD),再以评分标准为奖励进行 RL。在 HealthBench、ResearchQA 和 RubricHub Science 上,该框架得分高于所评估的其他后训练方法。

正文

View PDF HTML (experimental)

Abstract:Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level supervision, then as rewards for further RL. In the first stage, rubric-privileged on-policy distillation (RP-OPD), a student without access to the rubric matches a rubric-aware teacher's next-token distributions at student-generated prefixes. In the second stage, RL directly optimizes the rubric reward and improves beyond the observed distillation plateau. We evaluate the framework on health and science tasks using open-weight models. Across HealthBench, ResearchQA, and RubricHub Science, we compare post-training methods and vary the amount of SFT or RP-OPD training before RL, finding that our two-stage framework achieves the highest scores among the methods evaluated. RP-OPD + RL shows limited signs of reward hacking on RubricHub Science, whereas the SFT + RL baseline increasingly receives high rewards for claims of rubric compliance without providing the required content. These findings support using rubrics to guide on-policy distillation before applying rubric-based RL.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.02781 [cs.LG]
  (or arXiv:2610.02781v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02781

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xinpeng Wang [view email]
[v1] Fri, 2 Oct 2026 04:14:57 UTC (366 KB)

来源:arXiv:cs.LG · arxiv.org