arXiv:cs.AI· Johannes Wesch, Danni Liu, Jan Niehues·· 5 小时前AI 评分43
KV²:一种自精炼的 KV Cache 压缩方法
KV$^2$: A Self-Refining KV Cache
AI 导读
KV² 是一种查询无关的 KV cache 压缩方法,先用轻量代理打分器筛选信息量高的上下文 token,再仅对这一子集重算最终淘汰分数。在 RULER 16K 上以 2% KV cache 预算运行时,其平均分比次优基线高出 40 个百分点以上;在 LongBench 的 2%-10% 预算区间取得最高平均分,且压缩阶段运行时间和峰值内存均低于全上下文重建。代码已开源。
正文
Abstract:The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at this https URL.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.03198 [cs.AI] |
| (or arXiv:2610.03198v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03198 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Danni Liu [view email]
[v1]
Fri, 2 Oct 2026 12:11:35 UTC (1,107 KB)
来源:arXiv:cs.AI · arxiv.org