跳到正文
arXiv:cs.LG· Davis Wertheimer, Haochen Shen, Ahan Gupta, Derrick Liu, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang·· 4 小时前AI 评分43

自剪枝 Transformer:用 Universal Attention 实现极致 KV-Cache 压缩

A Self-Pruning Transformer: Extreme KV-Cache Compression with Universal Attention

AI 导读

研究者提出 Universal Attention,一种端到端可训练的注意力架构,其复合衰减机制可作为自适应剪枝准则,剔除对注意力计算贡献最小的 token,同时保留 RoPE 嵌入与 Softmax 注意力。该架构在自然语言与合成任务上实现 10× KV-Cache 压缩,下游性能优于 SOTA 基线和未剪枝 oracle;在 16k 长度下实现 25× 压缩并展现出更强的长上下文泛化能力。

正文

View PDF HTML (experimental)

Abstract:The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embeddings with alternative decay-based mechanisms, which can then be used to prune KV-cache during inference. However, these decay functions have limited expressivity, and in practice devolve into sliding-window-like eviction patterns. In this work, we propose a unifying framework for complementary and novel decay mechanisms, capturing complex key statistics and interactions while preserving expressive RoPE embeddings and Softmax attention. The resulting Universal Attention is a highly expressive and end-to-end trainable architecture, whose composite decay mechanism acts as a natural, $\textit{adaptive}$ pruning criterion, removing tokens that contribute least to attention computation. Experimentally, Universal Attention achieves state-of-the-art $10\times$ compression on natural language and synthetic task data, while $\textit{improving}$ downstream performance compared to both state-of-the-art baselines and unpruned oracles. It further demonstrates superior long-context generalization with unprecedented $25\times$ compression at length 16k.
Comments: 27 pages, 3 figures, 14 tables
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.09051 [cs.LG]
  (or arXiv:2610.09051v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09051

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ahan Gupta [view email]
[v1] Tue, 6 Oct 2026 20:00:04 UTC (1,592 KB)

来源:arXiv:cs.LG · arxiv.org