arXiv:cs.LG· Vladislav Gromadskii, David Li, Samson Gourevitch, Yazid Janati, Eric Moulines, Maxim Panov, Alexander Korotin·· 5 小时前AI 评分37
IDRF:掩码离散扩散模型的逆蒸馏奖励微调
IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
AI 导读
IDRF 是一种针对少步掩码离散扩散生成器的奖励微调框架,用逆蒸馏正则替代不可解的序列级 KL 惩罚,并证明其损失上界该 KL 散度。该方法无需参考模型 rollout,在 DNA、图像和文本生成上以最多少 32 倍的去噪步数取得高奖励,同时缓解 reward hacking 并保持样本质量。
正文
Abstract:Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03641 [cs.LG] |
| (or arXiv:2610.03641v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03641 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vladislav Gromadskii [view email]
[v1]
Fri, 2 Oct 2026 17:29:22 UTC (2,662 KB)
来源:arXiv:cs.LG · arxiv.org