跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu·· 15 小时前AI 评分45

统一分布训练框架 MGFlow:一步式视觉生成的新 SOTA

Unifying Distributional Training for One-Step Visual Generation

AI 导读

研究者提出统一理论框架 MGFlow,将分布建模与匹配差异分离,用高斯混合在全局矩与样本表示之间以可调粒度建模特征分布,同时支持最优传输与基于分数的匹配。在 ImageNet 256×256 上,MGFlow 大幅超越 FD-Loss 基线,pMF-H 上达到 1.45 FDr⁶、JiT-H 上 1.64 的 SOTA 结果。

正文

Authors:Chi Zhang, Shi Haoyang, Yueyi Liu, Ruichuan An, Junkang Zhou, Chang Li, Xiuyuan Lu, Yichi Zhang, Bo Wang, Yuhang Wu, Sen Cui, Miao Liu

View PDF HTML (experimental)

Abstract:Distributional training provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce a unified theoretical framework that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates MGFlow, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with 1.45 $\mathrm{FDr}^6$ on pMF-H and 1.64 on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore.
Comments: Project page: this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.35763 [cs.LG]
  (or arXiv:2609.35763v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.35763

arXiv-issued DOI via DataCite

Submission history

From: Haoyang Shi [view email]
[v1] Mon, 28 Sep 2026 17:59:14 UTC (34,829 KB)
[v2] Wed, 30 Sep 2026 17:47:11 UTC (35,680 KB)
[v3] Thu, 1 Oct 2026 17:58:34 UTC (40,216 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org