跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分47

SGLang 推出 Weight Cache Daemon:引擎重启权重加载从 495 秒降至 0.63 秒

Blog Fast Engine Recovery: Sub-Second Engine Restart for SGLang via Weight Cache Daemon Nowadays, State-of-the-Art (SOTA) models are getting much bigger and reloading the model service after a crash is very expensive. Therefore, we are introducing the Weight Cache Daemon, a persistent GP... Ant Ling Infra Team (Ant Group), Alibaba, SGLang Team August 21, 2026

AI 导读

蚂蚁集团基础设施团队联合阿里巴巴与 SGLang 团队推出 Weight Cache Daemon,通过 CUDA IPC 零拷贝映射将常驻 GPU 的量化后权重直接提供给新的 SGLang 引擎实例。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org