跳到正文
arXiv:cs.LG· Utkarsh Ranjan·· 3 小时前AI 评分39

WakeKV:面向会改变读取行为的注意力头的响应式可逆 KV 驻留策略

WakeKV: Reactive, Reversible KV Residency for Heads That Change Their Minds

AI 导读

WakeKV 提出一种响应式 KV 驻留策略,将"冷却"的注意力头状态移至可恢复的 CPU 存储池,而非冻结或永久驱逐。研究在 1.5B-8B 三个模型、三种场景下测量发现,多数注意力头在生成过程中至少改变一次读取行为。

正文

View PDF HTML (experimental)

Abstract:Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. Across three models (1.5B-8B) and three regimes (needle retrieval, long chain-of-thought, and multi-turn recall), we measure head behavior on four model-regime combinations and find that most heads change their reading behavior at least once during generation. We introduce WakeKV, a reactive residency policy that moves cooling heads to a recoverable CPU reservoir rather than freezing or permanently evicting their state. At matched memory or budget, WakeKV consistently improves miss rate over frozen classification and destructive eviction, evaluated across five model-regime combinations and over three cited baselines (SnapKV, uniform R-KV, and ReasonAlloc) across four eligible combinations. A FlexiCache/vLLM implementation on Mistral-7B confirms the benefit on real hardware, improving throughput while retaining LongBench quality.
Comments: Accepted to the NeurIPS 2026 Workshop on ML for Systems. 2 figures, 4 tables, appendix
Subjects: Computation and Language (cs.CL); Hardware Architecture (cs.AR); Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG)
Cite as: arXiv:2610.02713 [cs.CL]
  (or arXiv:2610.02713v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02713

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Utkarsh Ranjan [view email]
[v1] Fri, 2 Oct 2026 02:48:25 UTC (144 KB)

来源:arXiv:cs.LG · arxiv.org