跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分44

SGLang 用 Waterfill 与 LPLB 改进 DeepEP MoE 负载均衡

Blog Improving DeepEP MoE Load Balance in SGLang with Waterfill and LPLB Mixture-of-Experts (MoE) models rely on Expert Parallelism (EP) to scale inference across multiple GPUs. In SGLang, DeepEP and EPLB provide high-performance serving under EP, but the workload seen by ... NVIDIA Team June 26, 2026

AI 导读

SGLang 为 DeepEP MoE 推理引入 Waterfill 和 LPLB 两项调度期负载均衡特性。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org