跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 16 小时前AI 评分38

Ling-3.0-flash 在 Blackwell 上的 Batch-1 推测解码:TPOT 降至 0.78 ms

Blog Chasing the Batch-1 Floor: Ling-3.0-flash Speculative Decode on Blackwell Batch-1 decode keeps getting more important. Xiaomi MiMo, for example, announced MiMo-V2.5-Pro UltraSpeed in June, claiming 1,000 tok/s decode on a one-trillion-parameter MoE model. Batch 1 gives an ... RadixArk SGLang Team, Ant Ling Infra Team August 21, 2026

AI 导读

RadixArk SGLang 团队与 Ant Ling Infra 团队在 4 张 NVIDIA Blackwell GPU 上为混合线性注意力 MoE 模型 Ling-3.0-flash 优化 batch-1 解码,NEXTN/MTP 路径将单请求吞吐从 288 tok/s 提升至 606 tok/s,平均 TPOT 从 3.33 ms 降至 1.53 ms。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org