跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zhengyi Guo, Jiayuan Sheng, Wenpin Tang, David D. Yao·· 15 小时前AI 评分31

DOHF:基于 Doob's h-transform 引导的在线扩散模型微调

DOHF: Online Diffusion Fine-tuning with Doob's $h$-transform Guidance

AI 导读

研究者提出 DOHF(Diffusion Online h-guidance Fine-tuning),将 Doob's h-transform 转化为实用的在线训练算法,通过对生成样本分配最优性权重、估计归一化局部校正 ∇log h 并蒸馏进生成模型,支持黑盒与不可微奖励且无需额外网络评估。

正文

View PDF HTML (experimental)

Abstract:Reward-based diffusion fine-tuning faces practical challenges when desirable outcomes are rare or conditioning corrections are costly to estimate. In this work, we propose Diffusion Online $h$-guidance Fine-tuning (DOHF), which turns Doob's $h$-transform into a practical online training algorithm. DOHF assigns optimality weights to generated samples, estimates the normalized local correction $\nabla\log h$ under the current rollout policy, and distills it directly into the generative model. Theoretically, we characterize the population-optimal DiffusionNFT update as well as the various classfier free guidance methods through a unified $h$-transform perspective. Methodologically, our framework accommodates black-box and non-differentiable rewards without additional network evaluations. We further show improved alignments under three empirical scenarios. Our work demonstrates how adapting probabilistic conditioning through inexpensive estimation and iterative distillation can improve generative learning across statistical sampling and visual generation.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2609.31882 [cs.LG]
  (or arXiv:2609.31882v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.31882

arXiv-issued DOI via DataCite

Submission history

From: Zhengyi Guo [view email]
[v1] Fri, 25 Sep 2026 18:18:53 UTC (6,787 KB)
[v2] Thu, 1 Oct 2026 03:01:56 UTC (6,787 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org