跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel·· 14 小时前AI 评分36

iADD:改进扩散策略优化中的对齐与多样性

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

AI 导读

针对 DDPO 等扩散模型强化学习后训练方法在奖励优化时牺牲多样性与质量的问题,研究者提出 iADD,通过理论分析证明仅对扩散模型后段 timestep 更新可能损害多样性,并基于 Feynman-Kac 训练实现更优的对齐-多样性权衡。在三个任务上的实验与消融显示,该方法在对齐和多样性上均取得显著提升。

正文

View PDF HTML (experimental)

Abstract:Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.01789 [cs.LG]
  (or arXiv:2610.01789v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01789

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ashok Neupane [view email]
[v1] Thu, 1 Oct 2026 14:35:26 UTC (13,822 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org