LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分48
SGLang-JAX 在 TPU v7x 上优化 Ling-2.6-1T:用单个 Pallas Kernel 把 MoE 数据搬运藏在计算背后
Blog Optimizing Ling-2.6-1T on TPU with SGLang-JAX: Hiding MoE Data Movement Behind Compute with One Pallas Kernel SGLang-JAX now supports efficient serving of inclusionAI's Ling-2.6-1T on TPU v7x. With a working baseline in place, profiling pointed to the Mixture-of-Experts (MoE) path as the main bottleneck: each... Prayer, JamesBrianD, Haolin Fu, Haoguang Cai, Qinghan Chen June 17, 2026
AI 导读
SGLang-JAX 现已支持在 TPU v7x 上高效服务 inclusionAI 的 Ling-2.6-1T,其新 Pallas kernel Fused MoE V2 将 MoE prefill 延迟从 5.16 ms 降至 2.42 ms,降幅 53%。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org