arXiv:cs.LG· Yingxiang Yang, Weihang Xiao, Zhunxuan Wang, Joshua Flashner, Niresh Agarwal·· 4 小时前AI 评分39
BoT-GRPO:通过 Bag-of-Token 聚合实现高效过程奖励推理强化学习
BoT-GRPO: Efficient Process-Reward RL for Reasoning via Bag-of-Token Aggregation
AI 导读
BoT-GRPO 将 GRPO 扩展到 token 级奖励模型,通过长度无关的 bag of tokens 聚合计算每 token 优势,无需 critic、可直接替换 GRPO。
正文
Abstract:Reinforcement learning is now central to eliciting reasoning in large language models, while in the popular algorithm Group Relative Policy Optimization (GRPO) every token in a rollout receives the same advantage. We ask how to make process supervision efficient: accelerating convergence and improving final quality without the cost of value networks. We propose Bag-of-Tokens Group Relative Policy Optimization (BoT-GRPO), which extends GRPO to token-level reward models through a length-invariant "bag of tokens" aggregation: it collects all token-level rewards across rollouts, weights each by the inverse of its source sequence length, and computes per-token advantages relative to weighted group statistics. BoT-GRPO is critic-free, and is a drop-in replacement wherever GRPO is used when token-level reward is available. On React front-end code generation, BoT-GRPO reaches $80\%$ compile rate up to $1.9\times$ faster than GRPO and converges faster than modern GRPO variants (GSPO, DAPO, PURE) while reaching higher final compile and VLM-judged win rates. On a second task, AIME mathematical reasoning, BoT-GRPO delivers absolute Pass@$k$ gains up to $8.1\%$ over GRPO in half the steps. For both tasks we compare the algorithm's performance on reasoning vs. non-reasoning base-model families (Qwen2.5-3B, SmolLM3-3B, Phi-4-mini-reasoning). Our experiments also yield a practical recipe for the reward model itself: reward stability matters more than richness: clean, bounded, stable fine-grained signals consistently accelerate learning where noisier alternatives stall.
| Comments: | Published at the COLM 2026 Workshop on Efficient Reasoning |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09804 [cs.LG] |
| (or arXiv:2610.09804v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09804 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhunxuan Wang [view email]
[v1]
Wed, 7 Oct 2026 10:23:17 UTC (2,677 KB)
来源:arXiv:cs.LG · arxiv.org