跳到正文
arXiv:cs.AI· Yuxuan Lou, Kai Yang, Geng Zhang, Yong Liu, Yang You·· 3 小时前

DivMoE:通过跨领域专家组合实现细粒度 MoE Upcycling

DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition

AI 导读

DivMoE 提出细粒度 MoE Upcycling 框架,通过领域专用专家初始化和多样性约束路由,解决细粒度专家路由崩溃问题。在 Qwen3-1.7B 上,DivMoE 平均准确率达 55.6%,超越最强基线 51.6%;12B 参数模型经推理数据微调后以 64.5% 平均准确率追平 16B 的 Moonlight-MoE。

正文

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identify a structural pathology of fine-grained upcycling: when fine-grained experts are derived from a single source model, naive routing collapses and downstream accuracy drops to near-random (e.g., on Qwen3-1.7B, Drop-Upcycling-fine-grained reaches only 23.2% average accuracy across 15 benchmarks, essentially matching from-scratch training at 22.2%, while the same method's coarse-grained variant reaches 50.2%). We propose DivMoE, the first framework achieving fine-grained MoE Upcycling with structurally-balanced routing. DivMoE introduces domain-specialized fine-grained expert initialization, deriving experts from dense models that have undergone domain-adaptive continual pre-training, and diversity-constrained routing, a hard structural constraint guaranteeing that each token activates experts from distinct domain groups. Across two base models and 15 benchmarks, DivMoE consistently outperforms six upcycling baselines (55.6% vs. 51.6% for the strongest baseline on Qwen3-1.7B) and strictly improves over the dense base model on every benchmark after Stage 2 continual pre-training -- closing the regression gap that has plagued prior fine-grained upcycling. After supervised fine-tuning on a public reasoning mixture, our 12B-parameter DivMoE model matches Moonlight-MoE (16B) at 64.5% average accuracy while outperforming a controlled NVIDIA-Upcycling baseline by 6.1 percentage points.
Comments: 19 pages, published as a conference paper at NeurIPS 2026
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.11317 [cs.AI]
  (or arXiv:2610.11317v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11317

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kai Yang [view email]
[v1] Thu, 8 Oct 2026 06:23:01 UTC (1,434 KB)

来源:arXiv:cs.AI · arxiv.org