跳到正文
arXiv:cs.CL· Siddharth Bhandari, Lucas Gretta, Krishna Balasubramanian, Shiva Kasiviswanathan·· 3 小时前

ReadKV:面向 KV Cache 的查询自适应量化

Read What Matters: Query-Adaptive Quantization for KV Caches

AI 导读

ReadKV 提出一种查询自适应 KV cache 读取方法:键值以渐进式编码存储,每个查询按需分配 key-channel 与 value-token 前缀精度,存储内容保持不变。

正文

View PDF HTML (experimental)

Abstract:KV-cache entries are stored before their future queries are known, but each decoding query needs precision in different places. We study this mismatch using separate budgets for retained bits and bits fetched per query. ReadKV stores each key and value in a progressive code whose prefixes support different reconstruction precisions. For each query, it allocates key-channel prefixes using the query, computes attention from the reconstructed keys, and then allocates value-token prefixes using that attention. Stored entries remain unchanged. Each stage optimizes a calibrated distortion objective under a fixed budget; we prove exact allocation under diminishing refinement gains and relate these objectives to attention-output error. We also exhibit a finite-dimensional attention family where query-dependent access strictly outperforms every query-independent reader at the same read budget, even with unrestricted competing encoders and decoders.
Across six base models, reading four bits on average from an eight-bit cache increases C4 perplexity by at most 0.66%, using about one quarter of the logical reads and half the retained capacity of a 16-bit cache. It is consistently more accurate than storing and fully reading four bits at the same payload-read budget. Retaining more bits than each query fetches is aimed at long-context decoding, where the cache bytes moved per step, rather than the weights, dominate cost. Long-context question answering and retrieval on two instruction-tuned models provide additional quality evidence. On the tested 8K-token, batch-one, single-layer workload on an NVIDIA A10G, a restricted eight-bit ReadKV reader with a two-bit mean payload-read budget has 39% lower latency than the tested TurboQuant codec.
Comments: 50 pages, 3 figures, 12 tables
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Performance (cs.PF)
Cite as: arXiv:2610.11245 [cs.LG]
  (or arXiv:2610.11245v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11245

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Siddharth Bhandari [view email]
[v1] Thu, 8 Oct 2026 04:47:41 UTC (117 KB)

来源:arXiv:cs.CL · arxiv.org