arXiv:cs.LG(机器学习,全量分类)· Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang·· 14 小时前AI 评分40
CARM:面向 LLM 强化学习的取消感知响应掩码
CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
AI 导读
针对 LLM 强化学习中采样响应偏离策略的问题,研究者提出序列级掩码 CARM,通过对每个 token 对数概率比取绝对值再平均,避免正负漂移相互抵消。在数学推理任务上,CARM 在 AIME 2024/2025/2026 与 BeyondAIME 上的 mean@16 较几何平均掩码最高提升 3.13 个百分点;在四个代码基准上平均 pass@1 提升 2.88 个百分点。
正文
Abstract:Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make sampled responses off-policy. Sequence-level masking addresses this mismatch by deciding whether an entire response should contribute to optimization. A common masking rule uses the length-normalized geometric mean of sampled token probability ratios. Its signed log-ratios can cancel across positions, concealing substantial bidirectional policy drift. We propose \emph{Cancellation-Aware Response Masking} (CARM), a sequence-level mask that takes the absolute value of each token log-ratio before averaging, preventing opposing probability changes from canceling. We prove that accepted responses satisfy a joint bound on the fraction of sampled-token ratios outside a prescribed band and their mean log-distance beyond its boundaries. Experiments on mathematical reasoning and code generation show that CARM improves mean@16 averaged over AIME 2024/2025/2026 and BeyondAIME by up to $3.13$ percentage points over geometric-mean masking, and increases average pass@1 across four code benchmarks by $2.88$ points over the strongest evaluated baseline. These findings support CARM as a theoretically grounded and effective method for response-level off-policy control in LLM reinforcement learning.
| Comments: | 28 pages, 11 figures, 5 tables |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02039 [cs.LG] |
| (or arXiv:2610.02039v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02039 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yafei Zhang [view email]
[v1]
Thu, 1 Oct 2026 16:51:56 UTC (1,809 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org