跳到正文
arXiv:cs.CL· Zhiyun Shi·· 3 小时前

Attic-KV:只排练将被读取的内容,训练无关的 KV 缓存压缩方法

Rehearse Everything, Remember Nothing: Attic-KV Rehearses What Will Be Read

AI 导读

研究者提出 Attic-KV(Attic),一种训练无关的 KV 缓存压缩方法,用引用上下文的问答对让模型自我测验,替代通读全文的排练方式。

正文

View PDF HTML (experimental)

Abstract:Many key-value (KV) caches are compressed before anyone knows what will be asked of them: a document cached for retrieval, a prompt prefix shared across requests, the memory of a long conversation. The prevailing approach scores KV entries by rehearsal: the model rereads the context and keeps the entries it attends to, assuming that the more completely a cache rehearses its context, the better it remembers it. We show that under tight budgets this assumption backfires: rehearse everything, remember nothing. At a 3% keep ratio, rereading the whole context keeps 31.5 of 96.5 points on RULER, and on LongBench's natural-text tasks it falls below methods that rehearse nothing at all. The cause is that a cache keeps what it rehearses: rereading spreads the budget across the whole context, so the answer's own entries survive at little more than chance. Like a student before an exam, a cache remembers more by testing itself than by rereading. Two principles follow: rehearse what will be read, and rehearse as much as there is. We instantiate them as Attic-KV (Attic for short), a training-free rehearsal in which the model quizzes itself with question-answer pairs that quote the context, alongside anchor tokens in a content-adaptive amount. Changing only the rehearsal lifts three hosts that score it in three different ways: Attic alone is the best training-free method in all eight settings we test on RULER and LongBench's natural-text tasks, and plugged into the gradient-based KVgrad and the trained RestoreKV+, it raises them by up to 17.1 and 28.1 points. Its advantage grows as the budget shrinks, reaching 41.9 points over full rereading at a 3% keep ratio, and it compresses faster than rereading the whole context.
Comments: 14 pages, 5 figures
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.12133 [cs.CL]
  (or arXiv:2610.12133v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.12133

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhiyun Shi [view email]
[v1] Thu, 8 Oct 2026 15:22:50 UTC (844 KB)

来源:arXiv:cs.CL · arxiv.org