跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分44

SGLang 集成 Elastic EP:为 DeepSeek MoE 部署实现部分故障容错

Blog Elastic EP in SGLang: Achieving Partial Failure Tolerance for DeepSeek MoE Deployments To serve massive Mixture-of-Experts (MoE) models efficiently, deploying a "wide" Expert Parallelism (EP) strategy—often spanning 32 GPUs or more per inference instance—is not just an option; it is a n... The Mooncake Team, Volcano Engine March 25, 2026

AI 导读

SGLang 框架集成 Elastic EP,通过解耦专家与 GPU 的固定映射实现部分故障容错,DeepSeek V3.2 在 32 GPU、256 冗余专家配置下服务中断控制在 10 秒内,较完整重启的 2-3 分钟减少 90%。该方案静态性能与标准 DeepEP 持平,系统吞吐 3560.21 tokens/sec,由 Mooncake EP 提供容错通信层支持。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org