LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分47
Z Lab、Modal 与 SGLang 联合发布 DFlash 与 Spec V2 推测解码方案
Blog The next generation of speculative decoding: DFlash and Spec V2 Using Modal and Z Lab's DFlash speculative decoding models with SGLang’s newly default Spec V2 engine, you can achieve state-of-the-art latencies for LLM inference serving. Our new, jointly-released D... Z Lab, Modal, and SGLang Teams June 15, 2026
AI 导读
Z Lab、Modal 与 SGLang 联合发布 DFlash 推测解码模型及 SGLang 新默认 Spec V2 引擎,为 Qwen 3.5 397B-A17B 提供更低延迟的 LLM 推理服务。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org