跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Alessandro Montenegro, Riccardo Venturelli, Marco Mussi, Matteo Papini, Alberto Maria Metelli·· 14 小时前AI 评分32

PPO 中复用历史样本何时有效?wPPO-U 与 wPPO-BH 的系统研究

Reusing Past Samples in Proximal Policy Optimization: When and How Does It Help?

AI 导读

研究在多重重要性加权框架下实例化 wPPO-U 和 wPPO-BH 两种 PPO 变体,仅复用最近若干轮迭代的样本,以隔离数据复用本身的影响。两者保留 PPO 核心机制,分别采用普通重要性权重与 balance-heuristic 校正权重,并推导出策略改进下界为各自损失提供理论依据。作者在连续控制任务上实证考察数据复用何时、如何提升 PPO 的样本效率或最终性能。

正文

View PDF HTML (experimental)

Abstract:Among on-policy deep reinforcement learning methods, Proximal Policy Optimization (PPO) has become the de facto standard, due to its consistently strong empirical performance across diverse application domains. However, on-policy methods are inherently sample inefficient: fresh data collected under the current policy is used for just a few updates before being discarded. Off-policy methods avoid this inefficiency via experience replay, achieving notable sample efficiency gains, but at the cost of training instabilities or extensive tuning. This motivated the rise of hybrid strategies that augment PPO with off-policy data reuse. Existing sample-reuse variants of PPO demonstrated improved sample efficiency over vanilla PPO, yet a systematic study of when reuse helps, in which scenarios, and to what extent remains missing. In this work, we study the effectiveness of sample reuse in PPO by instantiating two variants within a multiple importance weighting framework. Both retain the core PPO mechanics, reusing only samples from a window of recent iterations, thereby isolating the effect of data reuse from other factors. The variants, termed wPPO-U and wPPO-BH, employ vanilla importance weights or balance-heuristic-corrected ones, respectively. For both, we derive policy improvement lower bounds providing theoretical grounding for their respective losses. We use them to empirically study when and how data reuse improves sample efficiency or final performance of PPO across continuous control tasks.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01399 [cs.LG]
  (or arXiv:2610.01399v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01399

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Alessandro Montenegro [view email]
[v1] Thu, 1 Oct 2026 10:05:35 UTC (5,180 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org