跳到正文
arXiv:cs.LG· Yongqiang Yao, Jinru Tan, Kaihuan Liang, Zixin Yin, Yazhe Niu, Ruihao Gong, Dahua Lin, Ningyi Xu·· 4 小时前AI 评分33

RollVerify:弥合长尾 rollout 强化学习的效率与精度

RollVerify: Bridging Efficiency and Accuracy in Long-Tail Rollout Reinforcement Learning

AI 导读

针对长尾 rollout 导致 GPU 空泡、异步部分 rollout 引入过期 off-policy 样本拖累精度的问题,研究者提出轻量级 RL 框架 RollVerify,在样本进入训练前主动验证与修复。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning is crucial for improving large language models' reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models' accuracy while preserving the efficiency gains of partial rollout. Experiments on mathematical and tool-assisted mathematical reasoning show that RollVerify achieves accuracy comparable to on-policy training while reducing training cost. Additional code-generation results provide preliminary evidence beyond mathematics.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09914 [cs.LG]
  (or arXiv:2610.09914v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09914

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yao Yongqiang [view email]
[v1] Wed, 7 Oct 2026 12:05:33 UTC (1,125 KB)

来源:arXiv:cs.LG · arxiv.org