arXiv:cs.LG· Sara Abdali, Jongwoo Ko, Pashmina Cameron·· 4 小时前AI 评分36
AttSVD:基于注意力引导 SVD 的提示词自适应低秩 KV Cache 压缩
AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD
AI 导读
AttSVD 提出一种免训练的低秩 KV Cache 压缩方法,按每条提示词自身的注意力几何在线做截断 SVD,只保留注意力实际读取的方向,并把每个注意力头的持久 KV 显存按保留秩成比例削减。
正文
Abstract:The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "interpretable" low-rank compression whose basis is derived from each prompt's own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes. Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06927 [cs.LG] |
| (or arXiv:2610.06927v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06927 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sara Abdali [view email]
[v1]
Sat, 3 Oct 2026 00:33:53 UTC (455 KB)
来源:arXiv:cs.LG · arxiv.org