arXiv:cs.LG· Davis Wertheimer, Haochen Shen, Ahan Gupta, Derrick Liu, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang·· 4 小时前AI 评分43
自剪枝 Transformer:用 Universal Attention 实现极致 KV-Cache 压缩
A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention
AI 导读
研究者提出 Universal Attention,一种端到端可训练的注意力架构,其复合衰减机制可作为自适应剪枝准则,剔除对注意力计算贡献最小的 token,同时保留 RoPE 嵌入与 Softmax 注意力。该架构在自然语言与合成任务上实现 10× KV-Cache 压缩,下游性能优于 SOTA 基线和未剪枝 oracle;在 16k 长度下实现 25× 压缩并展现出更强的长上下文泛化能力。
正文
Abstract:The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit{adaptive}$ pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $\textit{improving}$ downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented $25\times$ compression at length 16k.
| Comments: | 27 pages, 3 figures, 14 tables |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09051 [cs.LG] |
| (or arXiv:2610.09051v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09051 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ahan Gupta [view email]
[v1]
Tue, 6 Oct 2026 20:00:04 UTC (1,592 KB)
来源:arXiv:cs.LG · arxiv.org