arXiv:cs.CL· Zhiyun Shi·· 3 小时前
Attic-KV:只排练将被读取的内容,训练无关的 KV 缓存压缩方法
Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read
AI 导读
研究者提出 Attic-KV(Attic),一种训练无关的 KV 缓存压缩方法,用引用上下文的问答对让模型自我测验,替代通读全文的排练方式。
正文
Abstract:Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it. We show that under tight budgets this assumption backfires: rehearse everything, remember nothing. At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all. The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance. Like a student before an exam, a cache remembers more by testing itself than by rereading. Two principles follow: rehearse what will be read, and rehearse as much as there is. We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount. Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points. Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.
| Comments: | 14 pages, 5 figures |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12133 [cs.CL] |
| (or arXiv:2610.12133v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12133 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhiyun Shi [view email]
[v1]
Thu, 8 Oct 2026 15:22:50 UTC (844 KB)
来源:arXiv:cs.CL · arxiv.org