跳到正文
arXiv:cs.AI· Yue Zhang, Zun Wang, Han Lin, Yonatan Bitton, Idan Szpektor, Mohit Bansal·· 7 小时前AI 评分54

SPATIALUNCERTAIN 框架研究 VLM 能否判断何时不应回答空间问题

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?

AI 导读

论文提出 SPATIALUNCERTAIN 评估框架,研究视角相关的观测不确定性,覆盖遮挡导致证据缺失和透视导致证据误导两类失效模式。对八个开源与闭源视觉语言模型的测试显示,模型行为不随视觉证据可靠性变化,透视冲突时更倾向跟随投影外观而非物理 3D 关系;提示或微调未能完全解决,提供更好的视角比在同一误导观测上补充深度信息更有效。

正文

View PDF HTML (experimental)

Abstract:Spatial reasoning benchmarks typically evaluate whether vision-language models can derive the correct answer from a visual observation. Yet in real 3D environments, the observation itself may be unreliable: occlusion can remove task-relevant evidence, while perspective can make visible geometry misleading. Reliable spatial reasoning therefore requires more than answering a question correctly. A model must also assess whether its current observation provides sufficient and trustworthy evidence for that answer. We introduce SPATIALUNCERTAIN, a controlled evaluation framework for studying viewpoint-dependent observational uncertainty. We study two complementary failure modes: missing evidence caused by occlusion and misleading evidence caused by perspective. We further evaluate whether models can recognize when the current view is unreliable and identify a more informative observation. Across eight open- and closed-source vision-language models, we find that model behavior does not track the reliability of visual evidence. Models do not reliably become more cautious as evidence disappears, and under perspective conflict, their judgments increasingly follow projected appearance rather than the unchanged physical 3D relation. Internal analysis suggests a corresponding representational asymmetry: projected 2D relations are readily available, whereas the underlying physical 3D relation is barely decodable. Moreover, models that can identify an informative viewpoint when explicitly asked often fail to recognize when such an additional view is needed. These failures are not fully resolved by prompting or fine-tuning, and providing a better viewpoint is substantially more effective than adding depth information to the same misleading observation. Our results identify assessing the reliability of visual observations as a distinct and missing component of current spatial reasoning evaluation.
Comments: Website: this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2605.30557 [cs.CV]
  (or arXiv:2605.30557v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2605.30557

arXiv-issued DOI via DataCite

Submission history

From: Yue Zhang [view email]
[v1] Thu, 28 May 2026 20:44:47 UTC (13,319 KB)
[v2] Tue, 6 Oct 2026 01:04:05 UTC (14,230 KB)

来源:arXiv:cs.AI · arxiv.org