arXiv:cs.LG(机器学习,全量分类)· Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv·· 14 小时前AI 评分28
FERPO:前向熵正则化策略优化
FERPO: Forward Entropy-Regularized Policy Optimization
AI 导读
FERPO 是一种在线最大熵强化学习算法,仅用 critic 的价值预测做策略改进,无需对 critic 求动作梯度。它从熵与 KL 散度正则化的策略改进目标导出最优目标动作分布,再用自归一化重要性采样(SNIS)估计的前向 KL 目标拟合 actor。在 MuJoCo Playground 和 ManiSkill 上取得有竞争力的性能与样本效率提升,actor 更新速度也快于 REPPO。
正文
Abstract:Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).
| Comments: | Code: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.02198 [cs.LG] |
| (or arXiv:2610.02198v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02198 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sebastian Sanokowski [view email]
[v1]
Thu, 1 Oct 2026 17:59:41 UTC (774 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org