跳到正文
arXiv:cs.CL· Runguo Li·· 4 小时前AI 评分43

BreadthKV:长链式推理解码时 KV 缓存的精度—数量权衡

Spend Bytes on Breadth: Precision-Count Trade-offs for Decode-Time KV Compression in Long Chain-of-Thought Reasoning

AI 导读

BreadthKV 提出在固定字节预算下把更多字节用于缓存更多 token 而非更高精度,将量化与驱逐结合,并用 60 题端到端校准为每个模型和预算选择位宽。在三个推理模型和四个数学与科学基准上,它在 18 个设置中有 17 个优于单纯驱逐,并生成更短的输出;在最紧预算下 Qwen3-8B 的 AIME 样本触顶比例从驱逐的 91% 降至 40%。

正文

View PDF HTML (experimental)

Abstract:Reasoning models write most of their KV cache while decoding long chains of thought (CoT), so the cache has to be compressed online under a fixed memory budget. Decode-time methods mostly decide which tokens to evict. We ask how a fixed byte budget should be split between the number of cached tokens and their precision. BreadthKV spends the bytes on more tokens at low precision, combining quantization with eviction, and picks the bit-width for each model and budget with a 60-problem end-to-end calibration, since offline attention error does not predict it reliably. On three reasoning models and four math and science benchmarks, it scores above eviction alone in 17 of 18 settings and produces shorter outputs. Much of what eviction loses comes from derailed runs, which keep reasoning until the length cap without reaching an answer. On Qwen3-8B at our tightest budget, eviction sends 91% of AIME samples to the cap and BreadthKV 40%. Under the same protocol, BreadthKV is statistically indistinguishable from a joint rate-distortion allocator (RDKV) that uses 27% more KV memory-time, and it outperforms our re-implementation of ThinKV.
Comments: 17 pages, 4 figures, 13 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.05685 [cs.CL]
  (or arXiv:2610.05685v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.05685

arXiv-issued DOI via DataCite

Submission history

From: Runguo Li [view email]
[v1] Mon, 5 Oct 2026 01:56:50 UTC (113 KB)
[v2] Wed, 7 Oct 2026 04:10:06 UTC (117 KB)

来源:arXiv:cs.CL · arxiv.org