arXiv:cs.AI· Guanqun Zhao, Zijun Xie, Binbin Zheng, Yehan Yang, Jiafeng Lu, Aoqi Hu, Enlei Gong, Zeyu Chen·· 3 小时前
解构 Off-Policy 比率:面向异步强化学习、熵归一化的信任区域方法 ENTR
Deconstructing Off-Policy Ratios: Entropy-Normalized Trust Regions for Asynchronous Reinforcement Learning
AI 导读
针对异步强化学习中陈旧 off-policy 数据导致优化不稳定与策略崩溃的问题,研究者提出熵归一化信任区域方法 ENTR。该方法指出比率尺度由 token 熵决定,并识别出会放大训练—推理失配的低熵区间,仅按比率幅度设阈值会引入噪声、丢弃探索。
正文
Abstract:Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data destabilizes optimization and can cause policy collapse. Existing methods gate tokens by ratio magnitude alone, applying one threshold at every position. We show that the ratio's natural scale is set by token entropy, so deviations from mid-trajectory weight updates stay within this scale and carry genuine exploration. We further identify an overlooked low-entropy regime that breaks this scaling, where a near-zero probability amplifies train--inference mismatch into noise far beyond what the local entropy admits. A magnitude threshold admits this noise and discards the exploration. We therefore propose the Entropy-Normalized Trust Region (ENTR). Across long-horizon agentic tasks and mathematical reasoning benchmarks, ENTR outperforms existing asynchronous methods. It improves avg@1 on BrowseComp-Plus by $6.9\%$ over the strongest baseline, trains stably up to $30$ policy versions of staleness, and matches synchronous GRPO at a $2.6\times$ speedup.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22186 [cs.AI] |
| (or arXiv:2607.22186v5 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22186 arXiv-issued DOI via DataCite |
Submission history
From: Guanqun Zhao [view email]
[v1]
Fri, 24 Jul 2026 10:55:42 UTC (1,596 KB)
[v2]
Fri, 31 Jul 2026 03:34:43 UTC (1,594 KB)
[v3]
Mon, 3 Aug 2026 08:59:46 UTC (1,594 KB)
[v4]
Sat, 3 Oct 2026 13:19:19 UTC (1,818 KB)
[v5]
Thu, 8 Oct 2026 09:42:34 UTC (1,844 KB)
来源:arXiv:cs.AI · arxiv.org