跳到正文
arXiv:cs.LG· Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin·· 5 小时前AI 评分37

IDRF:掩码离散扩散模型的逆蒸馏奖励微调

IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models

AI 导读

IDRF 是一种针对少步掩码离散扩散生成器的奖励微调框架,用逆蒸馏正则替代不可解的序列级 KL 惩罚,并证明其损失上界该 KL 散度。该方法无需参考模型 rollout,在 DNA、图像和文本生成上以最多少 32 倍的去噪步数取得高奖励,同时缓解 reward hacking 并保持样本质量。

正文

View PDF HTML (experimental)

Abstract:Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03641 [cs.LG]
  (or arXiv:2610.03641v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03641

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Vladislav Gromadskii [view email]
[v1] Fri, 2 Oct 2026 17:29:22 UTC (2,662 KB)

来源:arXiv:cs.LG · arxiv.org