arXiv:cs.AI· Noam Elata, Itay Lamprecht, Mikey Shechter, Daniel Ohayon, Itay Hubara, Daniel Soudry·· 6 小时前AI 评分43
SAGA:解耦 key/value 头数的非对称稀疏注意力,LLM 长上下文解码提速超 2 倍
More Value per Key: Asymmetric Sparse Attention for Faster LLM Decoding
AI 导读
研究者提出 Sparse Asymmetric Group-Query Attention(SAGA),将 key 头与 value 头数量解耦,并搭配近似 top-N(Atop-N)稀疏注意力,在最长上下文下实现端到端解码速度超 2 倍于全注意力 GQA 基线,规模至 1.5B 参数。
正文
Abstract:Autoregressive generation in Large Language Models (LLMs) is constrained by the memory and computational demands of attention mechanisms. Sparse attention methods mitigate this cost by selecting only high-probability entries of the attention matrix. We observe that in many such methods, this renders the probability-value multiplication negligible, shifting the bottleneck to the query-key step. Key heads can therefore be reduced to accelerate inference, while retaining more value heads preserves capacity with limited additional decoding cost. We introduce Sparse Asymmetric Group-Query Attention (SAGA), which decouples key and value head counts to exploit this principle, and pair it with approximate top-N (Atop-N) attention, a simple sparse attention method designed to study the interaction between sparsity and head-count asymmetry. We formalize the benefits of this asymmetry theoretically and validate them empirically through latency measurements and quality evaluations on models up to 1.5B parameters. Together, SAGA and Atop-N achieve end-to-end decoding speedups exceeding $2\times$ over our full-attention GQA baseline at long contexts. Models trained from scratch with SAGA nearly match the quality of comparable GQA variants on the evaluated benchmarks. To facilitate adoption, we introduce an efficient fine-tuning method that converts pretrained models to the SAGA architecture, enabling practitioners to benefit from our approach without costly retraining.
| Comments: | Accepted to NeurIPS 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.04753 [cs.CL] |
| (or arXiv:2610.04753v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.04753 arXiv-issued DOI via DataCite |
Submission history
From: Noam Elata Mr [view email]
[v1]
Sat, 3 Oct 2026 20:38:43 UTC (275 KB)
[v2]
Tue, 6 Oct 2026 03:33:53 UTC (275 KB)
来源:arXiv:cs.AI · arxiv.org