arXiv:cs.LG· Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil·· 3 小时前AI 评分37
MFPG-PPO:多保真度策略梯度如何稳定数据稀缺的强化学习
Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
AI 导读
研究者将多保真度策略梯度(MFPG)框架扩展到现代 actor-critic 学习,提出 MFPG-PPO,通过重新设计采样、优势估计与控制变量构造来保持跨保真度相关性,并监控估计器不确定性以抑制方差膨胀。
正文
Abstract:Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develop MFPG for modern actor-critic learning in GPU-parallel simulation and on a physical robot. Our analysis and experiments show that naive extensions to proximal policy optimization (PPO) can lose cross-fidelity correlation or inflate variance. Our MFPG-PPO addresses these failures by redesigning the sampling, advantage estimation, and control variate construction to preserve cross-fidelity correlation, and by monitoring estimator uncertainty to prevent variance inflation. We also introduce a budget-aware MFPG-PPO to divide a fixed sampling budget among high- and low-fidelity data sources. Across simulated robot locomotion tasks of varying LF-to-HF transfer difficulty and HF data budgets, MFPG-PPO improves upon PPO trained on HF data alone in nearly all settings, and consistently matches the performance of PPO trained with 16x more HF data on the hardest task at the smallest HF budgets. In contrast, most baselines that use LF data perform well only where direct LF-to-HF transfer succeeds. MFPG-PPO enables stable learning on a physical Franka arm using only 4 real-robot episodes per update and no human demonstrations.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2610.02505 [cs.LG] |
| (or arXiv:2610.02505v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02505 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinjie Liu [view email]
[v1]
Thu, 1 Oct 2026 21:27:02 UTC (2,502 KB)
来源:arXiv:cs.LG · arxiv.org