arXiv:cs.LG· Vu Quang Hoang, Nghia Hieu Nguyen·· 3 小时前AI 评分26
稠密 MoE 作为重参数化宽 FFN:固定算力下的粒度扫描
Dense Mixture-of-Experts as a Reparameterized Wide FFN: A Granularity Sweep at Fixed Compute
AI 导读
研究用稠密类比隔离 MoE 动态专家组合的贡献:K 个全激活 SwiGLU 专家经 softmax 门控组合,固定 FFN 总宽度,K=1 为基线。验证损失随 K 非单调变化,K=2 比基线改善 0.0048,K=4 和 K=6 分别恶化 0.0053 和 0.0197。该架构等价于带 token 相关、单纯形约束组缩放的稠密 SwiGLU,单次运行结果显示固定宽度下 K=2 表现最佳。
正文
Abstract:Sparse Mixture-of-Experts (MoE) models combine learned routing with selection from a large expert pool. We isolate the contribution of dynamic expert combination using a dense analogue: $K$ SwiGLU experts, all active for every token and combined by a softmax gate, at fixed total FFN width. With no larger pool or discrete selection, the dense baseline is the $K=1$ case. Validation loss varies non-monotonically with $K$: $K=2$ improves over the baseline by $0.0048$, whereas $K=4$ and $K=6$ worsen it by $0.0053$ and $0.0197$, respectively. Routing generally remains soft and expert usage balanced, except in the first layer of the $K=4$ model, where routing is nearly one-hot. Forcing this gate to uniform increases loss by $2.4$ nats on a diagnostic subset, indicating that its concentrated routing is functionally important. We further show that the architecture is exactly a dense SwiGLU with token-dependent, simplex-constrained group scaling and contains the dense baseline in its function class. These single-run results suggest a granularity sweet spot for always-active, softly gated FFNs at fixed width, with $K=2$ performing best among the configurations tested.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02584 [cs.LG] |
| (or arXiv:2610.02584v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02584 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nghia Hieu Nguyen [view email]
[v1]
Thu, 1 Oct 2026 23:31:06 UTC (34 KB)
来源:arXiv:cs.LG · arxiv.org