跳到正文
arXiv:cs.CL· Neha Verma, Sungwon Kim, Kenton Murray, Kevin Duh·· 3 小时前

VFold:对称感知的跨层 Value Cache 压缩

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

AI 导读

研究者提出 VFold,一种对称感知的 value cache 合并策略,可在不修改 LLM 架构、不增加解码开销的前提下压缩 KV cache 内存并避免性能退化。该方法可与高比例量化或 key cache 剪枝叠加,达到单一方法无法实现的压缩比。

正文

View PDF HTML (experimental)

Abstract:While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this approach can be exploited alongside existing cache compression techniques, composing with high-ratio quantization or key cache pruning to reach compression ratios that neither method reaches alone, with minimal additional cost. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, offering a simple yet highly effective direction for scaling context windows under memory constraints.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.12338 [cs.CL]
  (or arXiv:2610.12338v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.12338

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Neha Verma [view email]
[v1] Thu, 8 Oct 2026 17:14:28 UTC (251 KB)

来源:arXiv:cs.CL · arxiv.org