跳到正文
arXiv:cs.LG· Chenliang Li, Adel Elmahdy, Alex Boyd, Zhongruo Wang, Siliang Zeng, Alfredo Garcia, Parminder Bhatia, Taha Kass-Hout, Cao Xiao, Mingyi Hong·· 7 小时前AI 评分33

SORL:用轮次级重要性采样与裁剪触发归一化稳定长程 LLM 智能体的离策略训练

Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalization

AI 导读

针对 PPO、GRPO 等 RL 算法在离策略训练多轮 LLM 智能体时易出现优化不稳定和性能崩溃的问题,研究者提出 SORL 框架,并实例化出 SO-PPO 与 SO-GRPO 两种算法。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning (RL) algorithms such as PPO and GRPO are widely used to train large language models (LLMs) for multi-turn agentic tasks. However, in off-policy training pipelines, these methods can exhibit unstable optimization dynamics and are prone to perfor- mance collapse. Through empirical analysis, we identify two fundamental sources of instability in this setting: (1) a granularity mismatch between token-level policy optimization and turn- structured interactions, and (2) high-variance and unreliable gradient updates induced by off- policy importance sampling and inaccurate advantage estimation. To address these challenges, we propose SORL, Stabilizing Off-Policy Reinforcement Learning for Long-Horizon Agent Train- ing. SORL introduces mechanisms that align policy optimization with the structure of multi- turn interactions and adaptively suppress unreliable off-policy updates, yielding more conserva- tive and robust learning dynamics. Within this framework, we instantiate two stabilized algo- rithms: SO-PPO and SO-GRPO. Both algorithms are designed to mitigate gradient variance and prevent optimization collapse without requiring careful early stopping or heuristic tuning. We evaluate SO-PPO and SO-GRPO on benchmarks spanning open-domain QA, multi-hop QA, and medical multiple-choice QA, and further assess their transfer to asynchronous RL for mathe- matical reasoning by training on DAPO-Math-17k and validating on AIME-2024. These results demonstrate that SORL provides a practical, scalable, and general framework for stabilizing re- inforcement learning in multi-turn LLM agent training and asynchronous RL for mathematical reasoning.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2511.20718 [cs.LG]
  (or arXiv:2511.20718v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2511.20718

arXiv-issued DOI via DataCite

Submission history

From: Chenliang Li [view email]
[v1] Tue, 25 Nov 2025 05:54:02 UTC (682 KB)
[v2] Tue, 24 Feb 2026 23:08:24 UTC (820 KB)
[v3] Mon, 5 Oct 2026 20:37:23 UTC (2,170 KB)

来源:arXiv:cs.LG · arxiv.org