arXiv:cs.AI· Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao·· 4 小时前AI 评分36
从后验集中现象重新思考基于概率的强化学习
Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
AI 导读
研究揭示基于概率奖励的无验证器强化学习存在长度依赖的失效模式——后验集中现象(PCP):推理链越长,参考答案的条件概率越坍缩到低方差区间,导致奖励几乎无法区分,使 GRPO 类策略优化不稳定。为此提出 RLCPR 框架,包含不确定性感知数据采样与集中感知正则化两个组件,在七个基准中的六个上比 SOTA 无验证器 RL 基线最高提升 4.0%,同时提升 token 效率。
正文
Abstract:Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a length-dependent failure mode of probability rewards, which we call the Posterior Concentration Phenomenon (PCP). We show that the probability of a reference answer conditioned on a reasoning trace often collapses to a low-variance interval as the trace becomes lengthy. This phenomenon results in nearly indistinguishable rewards, which, under GRPO-based settings, makes probability-based policy optimization unstable and inefficient. Motivated by this, we propose Reinforcement Learning with Concentration-aware Posterior Rewards (RLCPR), a verifier-free RL framework to explicitly account for PCP for better optimization stability and token efficiency. It has two components: uncertainty-aware data sampling, which reduces concentration-prone rollouts before generation, and concentration-aware regularization, which penalizes unnecessarily long traces when posterior rewards collapse. Extensive experiments show that, alongside higher token efficiency, RLCPR outperforms the state-of-the-art verifier-free RL baseline by up to 4.0% on six of seven benchmarks, including general-domain and mathematical reasoning challenges.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.01458 [cs.AI] |
| (or arXiv:2610.01458v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01458 arXiv-issued DOI via DataCite |
Submission history
From: Shiu-Hong Kao [view email]
[v1]
Thu, 1 Oct 2026 10:53:23 UTC (647 KB)
[v2]
Fri, 2 Oct 2026 04:54:33 UTC (647 KB)
来源:arXiv:cs.AI · arxiv.org