arXiv:cs.LG· Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim, Jae W. Lee·· 4 小时前AI 评分39
SketchSSM:全状态写入、紧凑草图读取的线性注意力状态访问优化
SketchSSM: Write to the Full State, Read from a Compact Sketch
AI 导读
SketchSSM 通过低秩状态加权查询近似,在保持全状态更新的同时用紧凑草图近似状态读取,把状态访问流量降低约 10 倍,并在四个 decode benchmark 上匹配 FP32 全状态基线的平均精度。
正文
Abstract:Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent-state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full state once to precompute outputs for these basis vectors, storing them in a compact sketch. Each subsequent decode step combines the sketch vectors with query-dependent coefficients to reconstruct the output without a full-state read. Across four Mamba-2-, GDN-, and KDA-based models, SketchSSM at a mean sketch rank of 8 reduces state-access traffic by approximately 10$\times$ while matching the average accuracy of the FP32 full-state baseline across four decode benchmarks, and preserves recall on four RULER retrieval tasks. At this rank on one NVIDIA B300, linear-attention kernel speedups over the Standard vLLM baseline reach 7.30$\times$, 5.02$\times$, and 5.24$\times$ for Mamba-2, GDN, and KDA, respectively, with up to 2.77$\times$ higher decode throughput on Nemotron 3 Super.
| Subjects: | Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC) |
| Cite as: | arXiv:2609.33051 [cs.LG] |
| (or arXiv:2609.33051v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.33051 arXiv-issued DOI via DataCite |
Submission history
From: Omin Kwon [view email]
[v1]
Sun, 27 Sep 2026 00:31:35 UTC (321 KB)
[v2]
Tue, 6 Oct 2026 03:34:01 UTC (435 KB)
来源:arXiv:cs.LG · arxiv.org