arXiv:cs.LG· Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong·· 7 小时前AI 评分33
SORL:用轮次级重要性采样与裁剪触发归一化稳定长程 LLM 智能体的离策略训练
Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization
AI 导读
针对 PPO、GRPO 等 RL 算法在离策略训练多轮 LLM 智能体时易出现优化不稳定和性能崩溃的问题,研究者提出 SORL 框架,并实例化出 SO-PPO 与 SO-GRPO 两种算法。
正文
Abstract:Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2511.20718 [cs.LG] |
| (or arXiv:2511.20718v3 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2511.20718 arXiv-issued DOI via DataCite |
Submission history
From: Chenliang Li [view email]
[v1]
Tue, 25 Nov 2025 05:54:02 UTC (682 KB)
[v2]
Tue, 24 Feb 2026 23:08:24 UTC (820 KB)
[v3]
Mon, 5 Oct 2026 20:37:23 UTC (2,170 KB)
来源:arXiv:cs.LG · arxiv.org