arXiv:cs.LG· Xin Teng, Muxiao Li, Hongyi Wen·· 4 小时前AI 评分44
分布鲁棒 MoE 训练:DRMoET 提升稀疏专家模型路由可靠性
Distributionally Robust Mixture-of-Experts Training
AI 导读
研究者提出 Distributionally Robust MoE Training(DRMoET),一种即插即用的训练目标,将逐层专家视为内生鲁棒性分组,按 EMA 平滑、激活加权的专家损失以熵正则化 softmax 规则更新每层专家分布,加强潜在非 top 路由路径。
正文
Abstract:Mixture-of-Experts (MoE) transformers scale capacity by activating only a few experts per token, but this sparsity creates a hidden reliability problem: when routing is imperfect, load-balanced models may send tokens to experts that are insufficiently trained for the assigned inputs. We propose Distributionally Robust MoE Training (DRMoET), a drop-in objective that treats layer-wise experts as endogenous robustness groups and optimizes high-loss routing outcomes rather than merely equalizing traffic. DRMoET updates a per-layer expert distribution by an entropy-regularized softmax rule on EMA-smoothed, activation-weighted expert losses, strengthening plausible non-top routing paths while preserving standard MoE computation. Under the FLAME-MoE recipe at 746M-total and 10.3B-total scales, DRMoET improves downstream averages over both standard FLAME-MoE and auxiliary-loss-free balancing. At 10.3B total parameters and 67B training tokens, DRMoET improves the seven-task average from 0.6625 to 0.6767, while the auxiliary-loss-free baseline achieves 0.6431. Mechanistic analyses show lower expert-loss variance with nearly unchanged mean loss, 4.3% lower excess loss under forced mid-$k$ misrouting, and improved domain-expert specialization. These results position routing robustness-not only utilization balance-as a practical objective for reliable sparse MoE scaling. Project page and code are available at: this https URL.
| Comments: | In proceedings of NeurIPS 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07207 [cs.LG] |
| (or arXiv:2610.07207v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07207 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xin Teng [view email]
[v1]
Mon, 5 Oct 2026 18:20:02 UTC (605 KB)
来源:arXiv:cs.LG · arxiv.org