跳到正文
arXiv:cs.LG· Ruwen Fan (Jimmy), Yuezhi Zu (Jimmy), Junru Li (Jimmy), Qingda Hu (Jimmy), Xinjun (Jimmy), Yang, Jiwu Shu, Youyou Lu·· 4 小时前AI 评分45

CoMoE:在消费级 GPU 上以主机为中心路由实现高效 MoE 推理

Democratizing MoE inference on commodity GPUs with CoMoE

AI 导读

CoMoE 是一套面向消费级 GPU 的通信高效 MoE 推理系统,通过以主机为中心的路由减少通信量并消除全局同步停顿。它采用主机托底的 token 多播与 token 级细粒度聚合,在 RTX 5090 上推理吞吐最高提升 1.46x。该方案以仅 23.4% 的硬件成本接近支持 NVLink 的 A800 GPU 性能,且无需 P2P 互联。

正文

View PDF HTML (experimental)

Abstract:Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks.
We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
Subjects: Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
Cite as: arXiv:2610.09424 [cs.DC]
  (or arXiv:2610.09424v1 [cs.DC] for this version)
  https://doi.org/10.48550/arXiv.2610.09424

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruwen Fan [view email]
[v1] Wed, 7 Oct 2026 04:24:45 UTC (549 KB)

来源:arXiv:cs.LG · arxiv.org