arXiv:cs.LG· Naoki Nishikawa, Taiji Suzuki·· 4 小时前AI 评分41
Transformer 强化学习分层推理奖励:达到 minimax 最优速率
Reinforcement Learning for Hierarchical Reasoning Rewards: Minimax-Optimal Rates with Transformers
AI 导读
研究证明,基于 Transformer 的 actor-critic 算法在分层奖励的推理任务后训练中可达到 minimax 最优速率(查询预算与正则化强度,至多对数因子),且在 prompt 数固定时同样最优。
正文
Abstract:Reinforcement learning (RL) has become a standard tool for post-training language models on reasoning tasks, where the policy is updated by reward feedback while exploring the space of responses. Despite its empirical success, theoretical understanding of RL post-training remains limited, in particular of why on-policy exploration combined with a neural reward model is effective. In this paper, we address this question by modeling the reward as a hierarchical function on the response space: the reward consists of infinitely many local components, each of which becomes relevant only after the preceding ones have been resolved. We show that a natural Transformer-based actor--critic algorithm, which alternates between sampling from the current KL-regularized policy, fitting a Transformer critic to the observed rewards, and updating the policy, achieves the minimax optimal rates in the query budget and in the regularization strength up to logarithmic factors, and is minimax optimal for a fixed number of prompts. In contrast, we prove that sampling from the fixed reference distribution, as in offline reward modeling, can limit regret decay to a logarithmic rate. These results show that on-policy exploration progressively zooms in on the region where the reward is concentrated, and quantify its benefit for RL post-training.
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.08561 [cs.LG] |
| (or arXiv:2610.08561v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08561 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Naoki Nishikawa [view email]
[v1]
Tue, 6 Oct 2026 15:43:08 UTC (212 KB)
来源:arXiv:cs.LG · arxiv.org