arXiv:cs.CL· Hexuan Deng, Zihao Yan, Xuebo Liu, Shuo Nie, Yue Wang, Chen Wang, Zhaohua Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Min Zhang·· 3 小时前
GRPODropout:用更少 rollout 提升在线强化学习效果
GRPODropout: Less is More for Online Reinforcement Learning Rollouts
AI 导读
针对 GRPO 等强化学习方法中策略熵塌缩导致采样多样性下降的问题,研究者提出 GRPODropout:在标准更新前选择性剔除少量高概率、正优势的 rollout,并对保留的优势重新中心化。该方法仅改变 rollout 使用方式,计算开销可忽略,准确率高于原始 GRPO 且 actor 熵更高,同时用于更新的 rollout 样本更少,体现了"less is more"。
正文
Authors:Hexuan Deng, Zihao Yan, Xuebo Liu, Shuo Nie, Yue Wang, Chen Wang, Zhaohua Zhang, Tianwen Jiang, Qiuyong Xiao, Jihong Zhang, Min Zhang
Abstract:Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss of sampling diversity weakens exploration and limits further improvement. Existing methods address this issue either through algorithm-level interventions, such as reward modification and entropy/KL regularization, or through token-level reweighting. We investigate a complementary perspective: entropy collapse can also be mitigated by changing which generated rollouts contribute to policy updates. Under the same sampling budget, not all rollouts contribute positively to an update, and selectively excluding some can improve learning. To address this, we propose GRPODropout: before the standard update, we use a simple strategy that selectively removes a small number of high-probability positive-advantage rollouts and recenters the retained advantages. To motivate this design, we develop a rollout-level theoretical analysis that guides method design and threshold selection. The method changes only rollout usage, and adds negligible computational overhead. Experiments show higher accuracy than original GRPO and higher actor entropy while using fewer rollout samples for updates, illustrating "less is more." This work provides insight into RL rollout usage: removing some rollouts can improve performance. Code is available at this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11854 [cs.LG] |
| (or arXiv:2610.11854v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11854 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hexuan Deng [view email]
[v1]
Thu, 8 Oct 2026 12:38:17 UTC (141 KB)
来源:arXiv:cs.CL · arxiv.org