跳到正文
arXiv:cs.CL· Jeonghyo Song, YoungJoon Yoo·· 3 小时前

SAGE:面向视觉语言解码器视觉定位的注意力汇聚感知引导强调方法

SAGE: Sink-Aware Guided Emphasis for Visual Grounding in Vision-Language Decoders

AI 导读

研究揭示视觉语言模型解码器存在分层注意力汇聚现象:首尾层会出现与提示无关、固定落在同一批图像区域上的 PIS(Prompt-Invariant Sinks),中间层才由提示驱动并支撑视觉-语言对齐。

正文

View PDF HTML (experimental)

Abstract:Recent large vision-language models (VLMs) pair a visual encoder with a large language model (LLM) and perform well on diverse image-text tasks, yet their reliability is often limited by decoder attention pathologies that suppress visual evidence and exacerbate hallucinations. In this paper, we revisit visual attention sinks and uncover a structured, layer-dependent behavior: across prompts, early and late decoder layers exhibit prompt-invariant attention collapse onto the same few image regions, which we term PIS (Prompt-Invariant Sinks), whereas mid layers become prompt-conditioned and drive vision-language alignment. This split suggests that treating sinks as a uniform effect is incomplete. Building on this insight, we propose SAGE (Sink-Aware Guided Emphasis), a lightweight intervention that steers decoder attention away from PIS and toward query-dependent regions of interest (ROIs) using token-aligned ROI masks derived from standard vision backbones such as CLIP, ViT, and DINOv3. Evaluated on diverse vision-encoder + decoder-only LLM VLM families, SAGE improves visual grounding, reduces hallucinations, and yields consistent gains across public downstream vision-language benchmarks, including fine-grained visual discrimination settings where localized evidence is crucial, when instantiated with backbone-derived ROI masks.
Comments: Accepted to EMNLP 2026 Findings
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL)
Cite as: arXiv:2610.11469 [cs.CV]
  (or arXiv:2610.11469v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.11469

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jeonghyo Song [view email]
[v1] Thu, 8 Oct 2026 08:18:58 UTC (6,453 KB)

来源:arXiv:cs.CL · arxiv.org