arXiv:cs.AI· Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei·· 6 小时前AI 评分35
RGPO:用自适应理由脚手架让大语言模型学会推理
Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
AI 导读
研究者提出 Rationale-Guided Policy Optimization(RGPO),将真实理由作为临时脚手架而非固定模仿目标,仅把更高奖励的模型自生成解回传到无引导设定中,从而在不要求 off-policy 数据匹配 RL 任务格式的前提下缓解奖励稀疏。
正文
Abstract:On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptively leverages ground-truth rationale information according to the model's current capability while preserving its freedom to explore. Rather than treating reference solutions as fixed imitation targets, RGPO uses them as temporary scaffolds: rationales help the model generate improved responses, after which only higher-reward, model-generated solutions are transferred back to the original unguided setting. This design allows training to exploit available ground-truth information without requiring off-policy data to follow the same format as the RL task. Across both language-only and vision-language reasoning settings, RGPO consistently improves performance over RLVR baselines, and ablation studies show that adaptive rationale guidance is a key contributor to these gains. These results suggest that RGPO offers a practical and general approach for reducing reward sparsity, stabilizing reinforcement learning, and improving reasoning performance in both text-only and multimodal models.
| Comments: | NeurIPS 2026 |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07342 [cs.AI] |
| (or arXiv:2610.07342v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07342 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hoang Phan [view email]
[v1]
Mon, 5 Oct 2026 20:15:12 UTC (2,495 KB)
来源:arXiv:cs.AI · arxiv.org