跳到正文
arXiv:cs.LG· Jinwoo Kim, Shraddha Barke·· 6 小时前AI 评分34

理解强化学习中的富化:RLVR 中提示与中间引导的梯度偏差分析

Understanding Enrichment in Reinforcement Learning

AI 导读

研究从数学上证明,RLVR 中省略或截断重要性权重修正,等价于隐式重新加权奖励,并将由此产生的梯度误差分解为尺度、旋转和方差三部分。作者提出基于序列蒙特卡洛(SMC)的权重修正机制,在稳定性与粒子序假设下,把标准修正方差随样本长度的指数累积降为加性累积。在 Qwen3-1.7B 于 OpenMathReasoning 稀疏区间上的微调中,该分析解释了富化(无论是否修正)如何帮助避免稀疏域崩溃。

正文

View PDF HTML (experimental)

Abstract:When rewards are sparse, reinforcement learning with verifiable rewards (RLVR) often uses hints or intermediate guidance to generate more successful rollouts. This enrichment biases policy-gradient updates unless corrected via importance weights, but existing methods omit correction or truncate importance weights in order to avoid the high variance of correction. It thus remains unclear what exactly is gained or lost in RLVR by correcting enriched rollouts. We show, mathematically, that omitted or truncated correction implicitly reweights the defined reward, and we decompose the resulting gradient error into scale, rotation, and variance to explain their distinct effects on learning. To make correction practical, we develop a novel sequential Monte Carlo (SMC) weight correction mechanism that, under stability and particle-order assumptions, tempers the exponential compounding of standard correction variance over the length of a sample to an additive accumulation. We then apply our analysis of enrichment to interpreting the results of a fine-tune of Qwen3-1.7B on a sparse band of OpenMathReasoning, establishing a concrete mechanism of how enrichment, both corrected and uncorrected, helps avoid collapse in sparse domains. Our main contribution is to understand, in general, how enrichment and correction can affect training, as opposed to claiming that either mode of operation is superior to unenriched RL.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02846 [cs.LG]
  (or arXiv:2610.02846v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02846

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jinwoo Kim [view email]
[v1] Fri, 2 Oct 2026 05:34:41 UTC (157 KB)

来源:arXiv:cs.LG · arxiv.org