跳到正文
arXiv:cs.LG· Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hao Wang, Alborz Geramifard·· 3 小时前AI 评分33

AvoKV-E:面向长推理的负载感知 KV Cache 淘汰策略

AvoKV-E: Payload-Aware KV Cache Eviction for Long Reasoning

AI 导读

AvoKV-E 是一种免训练的 KV cache 淘汰策略,针对长输出推理中缓存瓶颈从固定提示词转移到生成轨迹的问题,先延迟近期状态的淘汰资格,再按候选归一化读取压力、键冗余度和 value 负载潜力对条目排序。在匹配的活跃 KV 预算下,AvoKV-E 在不同模型和数据集上持平或超越冗余感知、复现式和思维自适应淘汰基线,缓存越紧张增益越大。

正文

Authors:Han Yu, Wenhui Zhu, Xiwen Chen, Zhipeng Wang, Hejian Sang, Han Shi, Menglin Zhou, Xuanzhao Dong, Minzhou Huang, Rui Cai, Hao Wang, Alborz Geramifard

View PDF HTML (experimental)

Abstract:Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat cached entries as routing objects, estimating whether an old key will still be read, will recur, or can be replaced. This routing-only view overlooks two effects: low-attention entries can carry large value payloads whose removal changes future predictions, and newly generated states can appear stale before later queries have had a chance to read them. We introduce AvoKV-E, a training-free eviction policy that first delays eligibility for recent states and then ranks eligible entries using candidate-normalized read pressure, key redundancy, and value-payload potential. According to empirical evaluation across different models and datasets, AvoKV-E matches or exceeds redundancy-aware, recurrence-based, and thought-adaptive eviction baselines at matched active-KV budgets, with its largest gains in the tightest-cache regime. Component and counterfactual analyses further connect these gains to delayed observation, payload-aware scoring, redundancy, and scale-robust normalization. Together, the results show that long-reasoning KV eviction should preserve not only keys that are likely to be read, but also the value payloads that sustain the reasoning trajectory.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03007 [cs.LG]
  (or arXiv:2610.03007v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03007

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Han Yu [view email]
[v1] Fri, 2 Oct 2026 08:37:57 UTC (125 KB)

来源:arXiv:cs.LG · arxiv.org