跳到正文
arXiv:cs.CL· Radha Gulhane, Quentin Anthony, Beren Millidge·· 3 小时前AI 评分39

MoE 预训练中的专家耦合:用相关性放置与 token 重排降低 All-to-All 开销

Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling

AI 导读

MoE 预训练中路由器早期就学会将 token 分配给相关专家,研究发现 top-2 下每层 0.8% 的专家对会被 42% 的 token 同时选中。作者据此提出相关专家放置和 token 重排两种方法,在 Megatron-LM 中将 all-to-all 时间降低 1.16-2.63X,端到端步时最多降低 1.41X,且不改变路由决策或专家参数。

正文

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.
Subjects: Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2610.09372 [cs.CL]
  (or arXiv:2610.09372v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.09372

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Quentin Anthony [view email]
[v1] Wed, 7 Oct 2026 03:26:24 UTC (4,516 KB)

来源:arXiv:cs.CL · arxiv.org