LMSYS:Blog(Chatbot Arena 团队)·· 12 小时前AI 评分50
LMSYS 博客:DeepSeek-V4-Pro 在 H20 GPU 上的极限推理部署
Blog Pushing the Limits of Serving DeepSeek-V4-Pro DeepSeek-V4-Pro is a 1.6-trillion-parameter Mixture-of-Experts (MoE) model released with both FP8 and FP4 weights. Models at this scale naturally benefit from accelerators such as NVIDIA Blackwell GPU... Tianyu Zhang, Yusong Gao, Yun Zhang August 19, 2026
AI 导读
LMSYS 博客给出 1.6 万亿参数 MoE 模型 DeepSeek-V4-Pro 在 H20 GPU 上的服务优化方案,单节点 H20-141GB 在 batch size 1 下达到 271 output tokens/s,与 B300 的 383.7 tokens/s 相比解码性能差距缩小至 1.42×。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org