LMSYS:Blog(Chatbot Arena 团队)·· 15 小时前精选AI 评分75
SGLang SSD Expert Pack 让 DeepSeek-V4-Flash 和 Kimi-K3 跑在消费级硬件上
Blog Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack SGLang brings the core idea of SSD-LLaMA to MoE inference: keep routed experts that do not fit in VRAM and host RAM on an NVMe SSD, load only the experts selected by the router, and use Expert Pack la... WiCi AI Team, SGLang Team August 29, 2026
AI 导读
LMSYS 的 SGLang 团队发布 SSD Expert Pack 路径,把装不进 VRAM 和内存的 MoE 路由专家放在 NVMe SSD 上,只加载 router 选中的专家,用 Expert Pack 布局、direct I/O、pinned 缓冲和 GPU 缓存把 SSD 容量变成可用的存储层。
推荐理由
原文给出在单张消费级 GPU 上运行超大 MoE 模型的完整方案和实测数据,读者可以评估这条 SSD 专家卸载路径是否适合自己的本地部署。
来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org