E-MoE:面向非因子化扩散语言模型的增强型混合专家方法
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
针对掩码扩散模型反向过程按位置因子化、少步生成质量受限的问题,研究者提出增强型混合专家方法 E-MoE,用 MoE 主干的路由决策构建离散共享隐变量上的混合分布,且不增加激活参数量。在合成多模态基准、二值化 MNIST 和 LM1B 上,E-MoE 相比因子化基线改善了少步生成效果。
Published on Sep 29
Authors:
,
,
,
Abstract
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.
View arXiv page View PDF Add to collection
Get this paper in your agent:
hf papers read 2609.37533
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper 0
No model linking this paper
Cite arxiv.org/abs/2609.37533 in a model README.md to link it from this page.
Datasets citing this paper 0
No dataset linking this paper
Cite arxiv.org/abs/2609.37533 in a dataset README.md to link it from this page.
Spaces citing this paper 0
No Space linking this paper
Cite arxiv.org/abs/2609.37533 in a Space README.md to link it from this page.
Collections including this paper 0
No Collection including this paper
Add this paper to a collection to link it from this page.
来源:HuggingFace Daily Papers(社区热门论文) · huggingface.co