arXiv:cs.AI· Seulgi Kim, Zhixiong Zhang, Xinwei Zhang, Jie Ling, Ronn Shaw·· 6 小时前AI 评分42
GeoPID:从几何视角分解与引导视觉语言模型中的视觉信息
GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
AI 导读
研究者提出免训练框架 GeoPID,从几何角度将视觉语言模型中的多模态信息分解为冗余、模态独有和协同三部分,在 22 个 VLM 和 14 个 benchmark 上验证了需视觉 grounding 的问题中正确预测具有更强的视觉独有成分。基于此,GeoPID 在推理时沿视觉独有子空间选择性放大视觉表示,无需更新模型参数即将视觉 grounding 能力平均相对准确率提升 7.63%。
正文
Abstract:While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63\%.
| Comments: | Under Review |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08401 [cs.CV] |
| (or arXiv:2610.08401v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08401 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Seulgi Kim [view email]
[v1]
Tue, 6 Oct 2026 14:14:25 UTC (17,459 KB)
来源:arXiv:cs.AI · arxiv.org