跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分46

SGLang 推出 Unified Radix Cache:用一棵基数树统一混合模型前缀缓存

Blog Unified Radix Cache: One Tree for Hybrid Model Prefix Caching Prefix caching reuses KV when requests share the same token prefix. Under full attention, once the KV for a shared prefix is computed, it remains valid as more tokens are appended. A later request wit... Zhangheng Huang, Ke Bao, Yi Zhang, Jialin Ouyang, Sicheng Pan August 11, 2026

AI 导读

SGLang 发布 Unified Radix Cache,用单一 token 基数树统一 full attention KV、滑动窗口 KV 与 Mamba 循环状态的复用边界,替代此前的缓存类矩阵。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org