arXiv:cs.LG· Yunho Choi, Jongwon Lim, Woojin Ahn, Minjae Oh, Jeonghoon Shim, Yohan Jo·· 2 天前AI 评分41
POISE:用模型内部状态做价值估计的强化学习算法,被 NeurIPS 2026 接收
Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal States
AI 导读
研究者提出 POISE(Policy Optimization with Internal State Value Estimation),把模型内部状态当作价值模型,用轻量探针读取前向传播中已算出的信号来预测基线,并与策略在线同步训练。
正文
Abstract:Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models rests on variance reduction, which requires both a reliable baseline and high prompt diversity within each training batch. This is especially difficult in multi-domain training for general reasoning models, where prompts from different tasks induce highly diverse gradient signals. Existing approaches fall short in different ways: GRPO estimates its baseline as the group mean over rollouts from the same prompt, so an accurate baseline leaves fewer distinct prompts in the batch, while PPO avoids this trade-off by training a policy scale critic, roughly doubling the cost of training. We introduce POISE (Policy Optimization with Internal State Value Estimation), a reinforcement learning algorithm that turns the model's internal states into a value model. A lightweight probe reads the signals already computed during the forward pass to predict the baseline, and is trained online alongside the policy. To preserve gradient unbiasedness, we introduce a cross-rollout construction that predicts each rollout's value from an independent rollout's internal states. On Qwen3-4B and OLMo3-7B-Instruct-DPO across a six-domain verifiable-reward corpus, POISE outperforms other RLVR baselines while achieving more stable training. Moreover, the probe matches a separate LLM-scale value model, generalizes to various tasks, and remains accurate as the policy scales. By leveraging the model's internal representations, POISE enables stable policy optimization.
| Comments: | Accepted to NeurIPS 2026; Project Page: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2605.07579 [cs.LG] |
| (or arXiv:2605.07579v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2605.07579 arXiv-issued DOI via DataCite |
Submission history
From: Jongwon Lim [view email]
[v1]
Fri, 8 May 2026 10:49:36 UTC (1,308 KB)
[v2]
Mon, 11 May 2026 03:09:39 UTC (1,308 KB)
[v3]
Thu, 1 Oct 2026 02:48:05 UTC (516 KB)
来源:arXiv:cs.LG · arxiv.org