跳到正文
arXiv:cs.AI· Arnab Kanti Tarafder, Jaume Guasch-Mart\'i, Gokcen Kestor, Jie Ren·· 7 小时前AI 评分43

RelayMoE:面向分布式 MoE 训练的内存高效专家路由

Memory-Efficient Expert Routing for Distributed MoE Training

AI 导读

RelayMoE 是一种基于环结构的 MoE 执行模型,通过让专家权重或 token 在环上流转并本地计算,避免构建完整的 top-k 扩展调度缓冲区。在 30B–57B 生产级 MoE 模型上,单层实验相较 Megatron-LM 平均提速 2 倍,全模型训练吞吐最高提升 2.02 倍,可训练序列长度最高扩展 2.85 倍。该模型还利用环结构在反向传播中实现内存高效的重计算。

正文

View PDF HTML (experimental)

Abstract:As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.
Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Networking and Internet Architecture (cs.NI)
Cite as: arXiv:2610.07333 [cs.DC]
  (or arXiv:2610.07333v1 [cs.DC] for this version)
  https://doi.org/10.48550/arXiv.2610.07333

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arnab Tarafder [view email]
[v1] Mon, 5 Oct 2026 20:07:27 UTC (8,974 KB)

来源:arXiv:cs.AI · arxiv.org