跳到正文
arXiv:cs.CL· Tianhao Chen, Yuheng Wu, Kelu Yao, Xiaogang Xu, Xiaobin Hu, Dongman Lee·· 4 小时前AI 评分34

BACON:面向多模态 KV Cache 压缩的边界注意力校准方法

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

AI 导读

针对多模态大语言模型长视觉上下文下 KV cache 过大、解码延迟高的问题,研究者提出即插即用方法 BACON,用 last query attention 校准 observation window attention,并通过层内一致性与层间持续性抑制噪声。

正文

View PDF HTML (experimental)

Abstract:Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evidence and discard answer-relevant tokens under aggressive compression. We identify last query attention as a complementary signal for recovering such evidence, though its irrelevant signals may introduce additional noise. We propose BACON, a plug-and-play method that calibrates observation window attention with last query evidence while suppressing noise through intra-layer coherence and inter-layer persistence. Across diverse benchmarks, models, budgets, and compression methods, BACON improves multimodal KV-cache compression by 7.5% on average under the most aggressive budget, with gains up to 30.9%.
Comments: EMNLP 2026 Oral
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as: arXiv:2606.14782 [cs.CV]
  (or arXiv:2606.14782v4 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2606.14782

arXiv-issued DOI via DataCite

Submission history

From: Tianhao Chen [view email]
[v1] Wed, 10 Jun 2026 10:09:59 UTC (7,550 KB)
[v2] Tue, 16 Jun 2026 10:51:58 UTC (7,546 KB)
[v3] Sun, 23 Aug 2026 09:01:01 UTC (7,551 KB)
[v4] Fri, 2 Oct 2026 07:05:35 UTC (7,551 KB)

来源:arXiv:cs.CL · arxiv.org