跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Qijia He, Ruinan Jin, Jun Luo, Shaofeng Zou, Yingbin Liang·· 14 小时前AI 评分39

异步 LLM 后训练:组质量封顶与收敛分析

Asynchronous LLM Post-Training: Group-Mass Capping and Convergence Analysis

AI 导读

研究提出 GMC-GRPO 方法,通过组质量封顶最小化加权估计器中的比率偏差,为异步 GRPO 类算法建立收敛保证。相比 TIC-GRPO,该方法将四阶延迟项的阈值依赖从 O(ε⁻⁴) 改善至 O(ε⁻²),调优步长后延迟相关项随组大小 G 以 G⁻²/⁵ 衰减。在 Qwen3 模型与推理基准上,GMC-GRPO 在大延迟下取得稳定基线中的最佳表现。

正文

View PDF HTML (experimental)

Abstract:Asynchronous reinforcement learning (RL) improves the efficiency of large language model post-training but introduces stale rollouts generated by earlier policies. Theoretical understanding of how this staleness affects convergence and how to mitigate its impact remains limited. We derive a convergence bound for GRPO-style algorithms that explicitly characterizes the tradeoff between the gradient estimator's second moment and bias. For trajectory-level importance-weighted estimators, our analysis shows that once the second moment is uniformly controlled, delay enters the bound through the bias introduced by clipping or rescaling. Guided by this insight, we propose a novel group mass capping GRPO (GMC-GRPO) method, which minimizes a ratio-based bias bound within a class of weighted estimators sharing a common second-moment guarantee. We establish convergence guarantees for asynchronous GMC-GRPO and show that, compared with TIC-GRPO, it improves the threshold dependence of the fourth-order delay term from $O(\epsilon^{-4})$ to $O(\epsilon^{-2})$ as $\epsilon\to0$, where $1+\epsilon$ is the ratio threshold. Under local policy overlap, the delay-dependent term decreases as $G^{-2/5}$ after tuning the step size, where $G$ is the group size. For fixed behavior and current policies, the bias introduced by group rescaling also vanishes as $G\to\infty$, whereas the bias from trajectory-wise clipping can persist. Experiments across Qwen3 models and reasoning benchmarks demonstrate improved robustness to stale rollouts, with GMC-GRPO achieving the best performance among stable baselines under large rollout delays.
Comments: 40 pages, 6 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01896 [cs.LG]
  (or arXiv:2610.01896v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01896

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Qijia He [view email]
[v1] Thu, 1 Oct 2026 15:45:08 UTC (1,162 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org