跳到正文
arXiv:cs.AI· Johannes Wesch, Danni Liu, Jan Niehues·· 5 小时前AI 评分43

KV²:一种自精炼的 KV Cache 压缩方法

KV$^2$: A Self-Refining KV Cache

AI 导读

KV² 是一种查询无关的 KV cache 压缩方法,先用轻量代理打分器筛选信息量高的上下文 token,再仅对这一子集重算最终淘汰分数。在 RULER 16K 上以 2% KV cache 预算运行时,其平均分比次优基线高出 40 个百分点以上;在 LongBench 的 2%-10% 预算区间取得最高平均分,且压缩阶段运行时间和峰值内存均低于全上下文重建。代码已开源。

正文

View PDF HTML (experimental)

Abstract:The memory footprint of the key-value (KV) cache constrains the practical use of long-context models, and it dominates cost when one prefilled context must later serve many different queries. In this reusable setting, query-agnostic compression trades cost against quality: lightweight estimators are cheap but less accurate, whereas full-context reconstruction scoring is more accurate yet reprocesses the entire prompt. We introduce KV$^2$, a query-agnostic KV-cache compression method based on selective reconstruction. KV$^2$ first uses a lightweight proxy scorer to identify informative in-context tokens, then reprocesses only this subset to compute final eviction scores. On RULER, Needle-in-a-Haystack, and LongBench, KV$^2$'s margin over baselines widens as the budget tightens: on RULER 16K at a 2% KV-cache budget it improves the average score over the next-best baseline by more than 40 percentage points, and on LongBench it attains the highest average across 2%-10% budgets at lower compression-stage runtime and peak memory than full-context reconstruction. Reusable KV-cache compression thus does not require reprocessing the full context. Our code is available at this https URL.
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.03198 [cs.AI]
  (or arXiv:2610.03198v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03198

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Danni Liu [view email]
[v1] Fri, 2 Oct 2026 12:11:35 UTC (1,107 KB)

来源:arXiv:cs.AI · arxiv.org