arXiv:cs.LG· Haeyong Kang, Chang D. Yoo·· 3 小时前AI 评分45
Budgeted Cache Repair:跨上下文 KV-Cache 复用的按预算缓存修复
Budgeted Cache Repair for Cross-Context KV-Cache Reuse
AI 导读
研究发现跨上下文 KV-Cache 复用并非无损:在 MMLU 和 GSM8K 上会带来明显精度损失,且按整次调用决定是否复用无法消除该损失。
正文
Abstract:Cross-context KV-cache reuse predicts a shared segment's keys and values under a new prefix instead of recomputing them, and has been reported to do so without quality loss. We find otherwise, and identify two problems. (1) A hidden cost: on MMLU and GSM8K, reuse costs substantial accuracy. (2) A decision at the wrong unit: no rule for deciding whether to reuse a cache removes that cost. What does help is choosing which parts of the cache to recompute, and the value of choosing well falls as the unit of choice grows: informed selection removes 49.5% of the cache error beyond chance at single rows (one token's keys and values), 10.6% at 64-token chunks, and nothing at the level of whole calls. Budgeted Cache Repair (BCR) acts at the unit where selection still pays. It drafts two tokens from the assembled cache, ranks cache rows by the attention those tokens pay them, and recomputes a fixed budget of rows exactly, in one of three layouts. The cost is paid rather than predicted away, and the draft that fails as a gate succeeds as a selector. BCR restores GSM8K to dense-prefill accuracy while still serving most calls from cache, and its best layout outperforms every reuse baseline's mean in the reference grid. The draft also beats a coin-flip selector at the same budget - a control prior evaluations lack.
| Subjects: | Hardware Architecture (cs.AR); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02233 [cs.AR] |
| (or arXiv:2610.02233v1 [cs.AR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02233 arXiv-issued DOI via DataCite |
Submission history
From: Haeyong Kang [view email]
[v1]
Sun, 27 Sep 2026 07:57:28 UTC (117 KB)
来源:arXiv:cs.LG · arxiv.org