arXiv:cs.LG· Zicheng Hu, Zhijian Zhou, Xuan Zhang, Yuchen Liu, Cheng Chen, Yuan Li, Qi Gu, Yan Feng, Hongyan Hao, Chao Qu·· 4 小时前AI 评分43
COPC:面向异步 LLM 强化学习的耦合离策略校正方法
COPC: Coupled Off-Policy Correction for Asynchronous LLM Reinforcement Learning
AI 导读
针对异步 RL 训练中策略与优势估计双重失配的问题,研究者提出 actor-critic 方法 COPC,将 token 级比率掩码与 TD 残差的双侧裁剪比率加权结合。在工具集成数学推理与搜索任务上,COPC 均超越最强异步基线,并在 64 步策略陈旧度下保持增益。相比同步 PPO 仍有 1.7 倍单步加速,附加开销极小。
正文
Abstract:Asynchronous RL accelerates large language model post-training by decoupling rollout generation from optimization, but trains on stale trajectories. Existing methods primarily correct token-level policy mismatch through importance-ratio control in the actor objective. We show that this \emph{policy-side correction} alone is insufficient: advantage estimates also inherit mismatch from behavior-policy continuations, which we term \emph{advantage staleness}. We derive exact bias and variance decompositions for a general two-channel actor update, revealing nonseparable coupling between policy-weight and advantage-estimation errors: their interaction induces multiplicative bias terms, while squared policy weights amplify advantage uncertainty in gradient variance. This motivates the hypothesis that policy- and advantage-side correction should be coordinated. We introduce Coupled Off-Policy Correction (COPC), an actor--critic method combining token-level ratio masking with two-sided clipped-ratio weighting of TD residuals for return and advantage estimation. Joint parameter sweeps across staleness levels support this hypothesis: the effect of one correction parameter depends on, and can reverse with, the other. COPC achieves the highest reported performance on tool-integrated mathematical reasoning and search, outperforming the strongest reported asynchronous baseline in each setting. It also offers a broad high-performing parameter region and improved training stability. In search, COPC remains stable throughout training, while most evaluated asynchronous baselines collapse late in training. These gains persist at 64-step policy staleness. COPC adds minimal step-time overhead over asynchronous PPO and retains a $1.7\times$ step-time speedup over synchronous PPO.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09597 [cs.LG] |
| (or arXiv:2610.09597v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09597 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zicheng Hu [view email]
[v1]
Wed, 7 Oct 2026 07:43:49 UTC (651 KB)
来源:arXiv:cs.LG · arxiv.org