跳到正文
arXiv:cs.CL· Kung-Hsiang Huang, Haoyi Qiu, Yutong Dai, Caiming Xiong, Chien-Sheng Wu·· 3 小时前

GUI-KV:面向 GUI 智能体的时空感知 KV 缓存压缩方法

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

AI 导读

GUI-KV 是一种免重训练的即插即用 KV 缓存压缩方法,用于加速基于视觉语言模型的 GUI 智能体。它结合空间显著性引导与时间冗余评分两项技术,在 AgentNetBench 的 5 张截图设置下将解码 FLOPs 降低 38.9%,同时比全缓存基线提升 4.1% 的步骤准确率。该工作已被 TMLR 接收。

正文

View PDF HTML (experimental)

Abstract:Graphical user interface (GUI) agents built on vision-language models have emerged as a promising approach to automate human-computer workflows. However, they also face the inefficiency challenge as they process long sequences of high-resolution screenshots and solving long-horizon tasks, making inference slow, costly and memory-bound. While key-value (KV) caching can mitigate this, storing the full cache is prohibitive for image-heavy contexts. Existing cache-compression methods are sub-optimal as they do not account for the spatial and temporal redundancy of GUIs. In this work, we first analyze attention patterns in GUI agent workloads and find that, unlike in natural images, attention sparsity is uniformly high across all transformer layers. This insight motivates a simple uniform budget allocation strategy, which we show empirically outperforms more complex layer-varying schemes. Building on this, we introduce GUI-KV, a plug-and-play KV cache compression method for GUI agents that requires no retraining. GUI-KV combines two novel techniques: (i) spatial saliency guidance, which augments attention scores with the L2 norm of hidden states to better preserve semantically important visual tokens, and (ii) temporal redundancy scoring, which projects previous frames' keys onto the current frame's key subspace to preferentially prune redundant history. Across standard GUI agent benchmarks and models, GUI-KV outperforms competitive KV compression baselines, closely matching full-cache accuracy at modest budgets. Notably, in a 5-screenshot setting on the AgentNetBench benchmark, GUI-KV reduces decoding FLOPs by 38.9% while increasing step accuracy by 4.1% over the full-cache baseline. These results demonstrate that exploiting GUI-specific redundancies enables efficient and reliable agent performance.
Comments: Accepted by TMLR
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2510.00536 [cs.CL]
  (or arXiv:2510.00536v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2510.00536

arXiv-issued DOI via DataCite

Submission history

From: Kung-Hsiang Huang [view email]
[v1] Wed, 1 Oct 2025 05:37:54 UTC (3,857 KB)
[v2] Wed, 7 Oct 2026 23:23:10 UTC (3,060 KB)

来源:arXiv:cs.CL · arxiv.org