arXiv:cs.LG· Yao Fu, Cyrus Chang, Ritchie Zhao, Bryce Long, Yueying Li, Mahdi Kamani, Samkit Jain, Rahul Raman, Tara Safavi, Shreya Gupta, Parsa Ashrafi Fashi, Minseok Lee, Julien Demouth, Bita Darvish Rouhani·· 4 小时前AI 评分37
SPIN:用于稀疏注意力的影子预测索引器
SPIN: Shadow Predictive Indexer for Sparse Attention
AI 导读
研究者提出 SPIN(Shadow Predictive Indexer),用基于历史的轻量预测识别重要 KV 块,避免每个解码步都为整个 KV cache 打分。在长上下文与智能体基准测试中,SPIN 实现 30-40% 稀疏度并保持任务质量;在 vLLM 端到端服务中,输出吞吐量最高提升 14.9%,中位 token 间延迟最多降低 13.2%。
正文
Authors:Yao Fu, Cyrus Chang, Ritchie Zhao, Bryce Long, Yueying Li, Mahdi Kamani, Samkit Jain, Rahul Raman, Tara Safavi, Shreya Gupta, Parsa Ashrafi Fashi, Minseok Lee, Julien Demouth, Bita Darvish Rouhani
Abstract:Indexer-based sparse attention reduces the cost of core attention by passing only a fixed, small number of important tokens to it. However, the indexer must still score the entire KV cache at every decoding step. This scoring overhead becomes a major bottleneck as the context length grows. We propose SPIN (Shadow Predictive Indexer) to reduce this indexer overhead. SPIN uses lightweight, history-based prediction to identify important KV blocks, avoiding the need to score the full KV cache at every decoding step. SPIN treats KV blocks and speculative decoding as first-class design and implementation considerations. Across extensive evaluations on long-context and agentic benchmarks, SPIN achieves 30-40% sparsity while preserving task quality. In end-to-end vLLM serving, SPIN improves output throughput by up to 14.9% and reduces median inter-token latency by up to 13.2%.
| Comments: | 12 pages, 4 figures |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09025 [cs.LG] |
| (or arXiv:2610.09025v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09025 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yao Fu [view email]
[v1]
Tue, 6 Oct 2026 19:23:10 UTC (1,033 KB)
来源:arXiv:cs.LG · arxiv.org