arXiv:cs.LG· Jiawei Lin, Saibo Geng, Thomas Bourgeat·· 2 天前AI 评分39
EchoPress:基于虚拟上下文重建的查询无关 KV Cache 剪枝
EchoPress: Query-Agnostic KV Cache Pruning via Virtual Context Reconstruction
AI 导读
EchoPress 是一种免训练的 KV cache 剪枝方法,利用标准 prefill 阶段已算出的 query 和 key 近似 KVzip 的重建注意力分数,每个请求只需重建第一个 chunk 来校准剩余上下文的重要性。
正文
Abstract:KV cache pruning reduces long-context inference memory usage by evicting less important key-value pairs. KVzip estimates importance through context reconstruction: prompting a model to repeat the context chunk by chunk. This achieves strong compression quality at the cost of additional forward passes. Learned approximations reduce this cost but require model-specific training. We analyze how KVzip identifies important cached information and show how to approximate its reconstruction scores using information already computed during prefill. These findings motivate EchoPress, a training-free method that approximates reconstruction attention using queries and keys from standard prefill. For each request, it reconstructs only the first chunk to calibrate importance scores for the remaining context. Experiments on LongBench and RULER with Qwen3-8B and Llama-3.1-8B-Instruct show that EchoPress matches KVzip in task accuracy across eviction ratios from 50% to 90%, while reducing compression overhead by a factor of 1.7-19.6 and total prefill time by a factor of up to 2.9. Code is available at this https URL.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00412 [cs.LG] |
| (or arXiv:2610.00412v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00412 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiawei Lin [view email]
[v1]
Wed, 30 Sep 2026 14:04:50 UTC (206 KB)
来源:arXiv:cs.LG · arxiv.org