arXiv:cs.CL· Hong Chen, Yudong Zeng, Yongwei Huang, Zuhao Ouyang, Dongnan Zheng, Junyan Zhang, Yubo Gao, Xuming Hu·· 7 小时前AI 评分40
TwinKV:微小的频率修正如何改变 KV Cache 压缩后的保留结果
Small Frequency Corrections Can Change What Survives KV Cache Compression
AI 导读
研究者提出免训练的 KV Cache 压缩方法 TwinKV,通过非局部的 post-RoPE key 频率对 value energy 打折,在精确存储预算下保留原始 key 和 value。
正文
Abstract:Compressing a key-value cache before its next question is known requires choosing what to retain without knowing which evidence will matter. Value energy measures entry strength but does not distinguish isolated keys from those with many similar neighbors. We introduce TwinKV, a training-free method that discounts value energy by nonlocal post-RoPE key frequency. Prefix attention allocates head capacities, while retained entries preserve their original keys and values under an exact storage budget. Across four language models, TwinKV exceeds five evaluated compressed baselines in mean score on LongBench, LooGLE, and RULER at 50\% KV removal. Component controls isolate the frequency contribution. On Llama-3.2-1B RULER at 75\% removal, normalized frequency weights average 0.95, yet change 7\% of nonprotected retained positions and improve value-only retention by about 5.5 points under both uniform and adapted capacities. Permuting the weights within heads weakens this gain. These results show that modest frequency corrections can change retention and answering outcomes, with effects that depend on the model and task.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2608.27128 [cs.CL] |
| (or arXiv:2608.27128v4 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.27128 arXiv-issued DOI via DataCite |
Submission history
From: Hong Chen [view email]
[v1]
Thu, 27 Aug 2026 13:43:30 UTC (86 KB)
[v2]
Mon, 31 Aug 2026 03:59:21 UTC (86 KB)
[v3]
Sat, 26 Sep 2026 14:22:17 UTC (187 KB)
[v4]
Tue, 6 Oct 2026 08:16:19 UTC (187 KB)
来源:arXiv:cs.CL · arxiv.org