跳到正文
arXiv:cs.LG· Qiyuan Huang, Tianshi Xu, Meng Li·· 3 小时前AI 评分50

如何理解并防止 RLVR 中的灾难性策略坍缩:GRPO 后期崩溃的机制与 Mesh Learning 方案

All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR

AI 导读

针对 LLM 的 RLVR 后训练中 GRPO 类算法出现的后期灾难性坍缩,研究提出统一理论框架,证明主流 RLVR 目标会逐步把概率质量集中到单一策略,而维持任务准确率需要最低策略容量,二者冲突即坍缩机制,并据此导出轻量在线预警信号 MEI。

正文

View PDF HTML (experimental)

Abstract:During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy capacity. The conflict between these two results provides a mechanistic explanation for catastrophic collapse. We further derive the {Mirrored Entanglement Index (MEI)} as a lightweight online warning signal. To prevent collapse, we propose \textbf{Mesh Learning}, which exposes multiple reasoning strategies and prevents any single strategy from dominating optimization. Across AIME26, AIME25, MATH-500, GPQA, and LiveCodeBench, Mesh Learning consistently outperforms strong baselines across Qwen and Phi model families, with gains of up to 13.4 pp and 11.5 pp, respectively. These results establish strategy preservation as a key principle for stable RLVR. Code is available at this https URL.
Comments: 84 pages, 10 figures
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02835 [cs.LG]
  (or arXiv:2610.02835v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02835

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Qiyuan Huang [view email]
[v1] Fri, 2 Oct 2026 05:27:33 UTC (14,395 KB)

来源:arXiv:cs.LG · arxiv.org