跳到正文
arXiv:cs.LG· Vu Quang Hoang, Nghia Hieu Nguyen·· 3 小时前AI 评分26

稠密 MoE 作为重参数化宽 FFN:固定算力下的粒度扫描

Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute

AI 导读

研究用稠密类比隔离 MoE 动态专家组合的贡献:K 个全激活 SwiGLU 专家经 softmax 门控组合,固定 FFN 总宽度,K=1 为基线。验证损失随 K 非单调变化,K=2 比基线改善 0.0048,K=4 和 K=6 分别恶化 0.0053 和 0.0197。该架构等价于带 token 相关、单纯形约束组缩放的稠密 SwiGLU,单次运行结果显示固定宽度下 K=2 表现最佳。

正文

View PDF HTML (experimental)

Abstract:Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02584 [cs.LG]
  (or arXiv:2610.02584v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02584

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nghia Hieu Nguyen [view email]
[v1] Thu, 1 Oct 2026 23:31:06 UTC (34 KB)

来源:arXiv:cs.LG · arxiv.org