arXiv:cs.AI· Zhicheng Cai, Xinyuan Guo, Hanlin Wu, Wei-Ying Ma, Ya-Qin Zhang, Hao Zhou·· 3 小时前
RIPO:用黎曼等距策略优化解决 LLM RL 中的探索崩溃
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
AI 导读
针对 PPO-Clip 因用欧氏度量衡量策略差异、与策略黎曼流形内在几何不一致而导致的探索崩溃问题,研究者提出黎曼等距策略优化(RIPO),在黎曼流形上保证等距策略更新以平衡探索与利用。RIPO 取得更优的偏差-方差权衡,在七个竞赛级基准上显著超越现有 LLM RL 算法,AIME24 上较 GRPO 最高提升 60%。
正文
Abstract:Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip are inherently limited by exploration collapse. Subsequent works remain primarily heuristic and fail to identify the essential cause of PPO-Clip's failure. This work reveals the fundamental flaw of PPO-Clip: it implicitly measures policy discrepancy using Euclidean metric, which is theoretically inconsistent with the intrinsic geometry on the policy Riemannian manifold. This geometric mismatch results in overly conservative updates in low-probability regions while aggressive in high-probability regions, ultimately collapsing exploration. To correct this geometric flaw, we propose Riemannian Isometric Policy Optimization (RIPO), which guarantees isometric policy updates on the Riemannian manifold, effectively balancing exploration and exploitation. We further show that RIPO achieves a favorable bias-variance trade-off, which stabilizes optimization. Extensive experiments demonstrate that RIPO significantly surpasses existing LLM RL algorithms across seven competition-level benchmarks (up to 60% improvement over GRPO on AIME24).
| Comments: | ICML 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.10169 [cs.LG] |
| (or arXiv:2607.10169v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2607.10169 arXiv-issued DOI via DataCite |
Submission history
From: Zhicheng Cai [view email]
[v1]
Sat, 11 Jul 2026 07:19:59 UTC (2,494 KB)
[v2]
Thu, 8 Oct 2026 16:00:32 UTC (2,493 KB)
来源:arXiv:cs.AI · arxiv.org