跳到正文
arXiv:cs.LG· Inbasekaran S·· 3 小时前AI 评分41

Page-EntroKV:面向分组查询注意力的硬件对齐熵加权 KV-Cache 淘汰框架

Page-EntroKV: Hardware-Aligned, Entropy-Weighted KV-Cache Eviction under Grouped-Query Attention

AI 导读

Page-EntroKV 提出一种在 GQA 实际分配粒度上运行的 KV-Cache 淘汰框架,用 sink 隔离的 Renyi-2 熵对组内注意力头加权池化,并将分数投影到 PagedAttention 页帧上按(层、组、页)执行淘汰。

正文

View PDF HTML (experimental)

Abstract:Serving long-context autoregressive language models is constrained by the key-value (KV) cache. Most dynamic eviction methods score token importance per query head and choose tokens independently. This fits poorly with grouped-query attention (GQA), where several query heads share one physical KV buffer: divergent per-head selections force the serving engine to retain the union of their choices - inflating the cache by up to the group ratio r - while arithmetic mean pooling dilutes the specialized retrieval heads that carry factual recall. We introduce Page-EntroKV, a formal framework for KV-cache eviction operating at the granularity GQA serving actually allocates. Heads within each physical group are pooled by weights derived from sink-isolated collision (Renyi-2) entropy - one inner product per head, computed once at prefill with no calibration - so sink heads cannot masquerade as retrieval heads. Pooled scores are projected onto PagedAttention page frames, and eviction executes at the hardware tuple (layer, group, page). We formalize the union overhead ratio (UOR) and intra-group disagreement, prove an exact identity linking them for two-head groups alongside two-sided bounds at every group ratio, prove strict budget preservation and a finite-context needle-retention bound that arithmetic mean pooling provably violates, and give exact per-layer page accounting. On a pilot architecture (Qwen2.5-1.5B-Instruct, r=6), head-independent replay over 2,240 group measurements yields union overhead up to 4.75x at a 2% budget, while Page-EntroKV holds UOR exactly 1.000; sink isolation removes a 13x sink masquerade; needle recall is 100% versus 0% for mean pooling at a 20% budget; retained cardinality is exact for every page size; and QA and code tasks remain solvable at 20% retention.
Comments: 24 pages, 8 figures, 10 tables. Formal framework with pilot-scale empirical validation on Qwen2.5-1.5B-Instruct. Includes step-by-step derivations (Appendix C) and PyTorch reference implementation (Appendix D). Code and data available at: this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.03135 [cs.LG]
  (or arXiv:2610.03135v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03135

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Inbasekaran S [view email]
[v1] Fri, 2 Oct 2026 11:01:39 UTC (1,814 KB)

来源:arXiv:cs.LG · arxiv.org