跳到正文
arXiv:cs.LG· Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang·· 3 小时前AI 评分37

奖励膨胀:强化学习的一种健康刺激

Reward Inflation: A Healthy Stimulus for Reinforcement Learning

AI 导读

研究者提出"奖励膨胀"(reward inflation),即在训练过程中逐步放大奖励幅度,理论证明其能引入隐式近因加权、加快策略适应,并通过维持梯度信号抑制休眠神经元、保持可塑性。在 ALE 游戏和 MuJoCo 任务上的实验显示,适度奖励膨胀对广泛任务有益。团队还提出自适应变体 Fed,可动态调整膨胀水平,效果常优于固定膨胀。

正文

View PDF HTML (experimental)

Abstract:Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these findings, showing that an appropriate level of reward inflation benefits a broad range of tasks. Finally, we introduce Fed, an adaptive variant that adjusts the inflation level on the fly, and find that it often improves upon fixed inflation.
Comments: Accepted at NeurIPS 2026
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02545 [cs.LG]
  (or arXiv:2610.02545v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02545

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ganghun Lee [view email]
[v1] Thu, 1 Oct 2026 22:33:15 UTC (5,694 KB)

来源:arXiv:cs.LG · arxiv.org