跳到正文
arXiv:cs.LG· Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham·· 4 小时前AI 评分49

视觉语言模型如何进行空间变量绑定:两项并行机制

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

AI 导读

视觉语言模型(VLMs)依赖两种并行机制来表示空间变量绑定。语言模型主干中的中间层在物体对应的视觉 token 上表示与内容无关的空间关系,但只起次要作用;空间信息的主要来源是视觉编码器,其表示编码了物体布局并被语言模型主干直接利用。

正文

View PDF HTML (experimental)

Abstract:Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechanisms to represent spatial variable binding. In the language model backbone, intermediate layers represent content-independent spatial relations on top of visual tokens corresponding to objects. However, this mechanism plays only a secondary role in shaping model predictions. Instead, the dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone. Notably, this spatial signal is distributed globally across visual tokens, extending beyond object regions into surrounding background areas. We validate the generalization of our findings to complex natural images from the COCO dataset, where globally amplifying the vision-derived spatial representations across all image tokens corrects spatial variable binding failures across models of various sizes. Together, our results clarify how spatial variable binding is computed within VLMs and highlight the central role of vision encoders in enabling it.
Comments: 66 pages, 81 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2603.22278 [cs.CV]
  (or arXiv:2603.22278v3 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2603.22278

arXiv-issued DOI via DataCite

Submission history

From: Nikhil Prakash [view email]
[v1] Mon, 23 Mar 2026 17:58:02 UTC (18,214 KB)
[v2] Thu, 4 Jun 2026 18:56:47 UTC (25,171 KB)
[v3] Tue, 6 Oct 2026 15:42:39 UTC (13,791 KB)

来源:arXiv:cs.LG · arxiv.org