跳到正文
arXiv:cs.CL· Saif Punjwani, Micah Goldblum·· 3 小时前AI 评分47

ExpDis:将探索与优化解耦的 RLVR 框架

Decoupling Exploration from Optimization in RLVR

AI 导读

研究者提出 Exploration-Distillation(ExpDis)框架,将 RLVR 中的探索与优化解耦:先用带新颖性奖励的 explorer 策略采样,过滤轨迹的正确性与质量后蒸馏到独立 student 策略,student 再以无新颖性奖励训练,并多轮交替进行。

正文

View PDF HTML (experimental)

Abstract:Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectories for correctness and quality, and distill them into a separate student policy. The student policy is then trained without a novelty bonus. We repeat the above procedure for several rounds, alternating between exploration and optimization. This decoupling allows us to aggressively scale exploration without degrading the student policy. Across seven mathematical reasoning benchmarks and two model families, ExpDis outperforms DAPO at the same wall-clock budget. Moreover, we observe improved pass@$k$ scaling, indicating that ExpDis produces models that generate more diverse correct solutions.
Comments: 20 pages, 16 figures, 9 tables. Code: this https URL. Checkpoints: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.10536 [cs.LG]
  (or arXiv:2610.10536v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10536

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Saif Punjwani [view email]
[v1] Wed, 7 Oct 2026 17:59:26 UTC (782 KB)

来源:arXiv:cs.CL · arxiv.org