跳到正文
原文
LMSYS:Blog(Chatbot Arena 团队)·· 13 小时前AI 评分46

HiSparse:用分层内存为稀疏注意力提速,GLM-5.1-FP8 长上下文吞吐最高提升 5 倍

Blog HiSparse: Turbocharging Sparse Attention with Hierarchical Memory Self-attention has become a major bottleneck in scaling LLMs to long contexts because of its quadratic compute and memory/IO cost. This has driven growing interest in efficient attention mechanisms. A... Zhiqiang Xie, Zhangheng Huang, Tingwei Huang April 10, 2026

AI 导读

LMSYS 团队提出 HiSparse 分层内存系统,将不活跃的 KV cache 卸载到主机内存、在 GPU HBM 保留热缓冲区,缓解稀疏注意力的显存容量瓶颈。在 GLM-5.1-FP8 上,8×H200 部署下 256 并发请求吞吐达基线 3 倍以上,长上下文场景最高提升 5 倍。目前支持 DeepSeek-V3.2 与 GLM-5.1 等采用 DSA 的模型,属实验性功能。

来源:LMSYS:Blog(Chatbot Arena 团队) · lmsys.org