arXiv:cs.LG(机器学习,全量分类)· Junhyun Ha, Juho Lee, Byoungwoo Park·· 17 小时前AI 评分33
PReFlow:用提案条件精炼流改进扩散策略
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
AI 导读
研究者提出 Proposal-Conditioned Refinement Flows(PReFlow),一种结合 critic 提案选择与条件精炼流的策略提取方法,通过 KL 正则化目标诱导高斯平滑行为先验下的 Gibbs 策略。
正文
Abstract:Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.
| Comments: | 27 pages, 10 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2609.36812 [cs.LG] |
| (or arXiv:2609.36812v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.36812 arXiv-issued DOI via DataCite |
Submission history
From: Junhyun Ha [view email]
[v1]
Tue, 29 Sep 2026 06:25:13 UTC (1,278 KB)
[v2]
Thu, 1 Oct 2026 17:05:17 UTC (1,278 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org