arXiv:cs.AI· Gouki Minegishi, Hiroki Furuta, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo·· 6 小时前AI 评分49
VLM 抽象推理内幕:匹配的是物体还是关系?机制分析揭示两条竞争回路
Matching Object or Relation? Tracing Abstract Reasoning Inside VLMs
AI 导读
研究采用心理学 Relational Match-to-Sample(RMTS)范式并结合机制分析,在 GPT、Claude、Gemini 等前沿 API 模型及 Qwen3.5、Gemma-4、InternVL3 三个开源家族上评测,识别出能力层级、模型规模、场景物体数量与逐物体刺激噪声四项影响 VLM 关系匹配的杠杆。
正文
Abstract:Vision Language Models (VLMs) excel on visual benchmarks but fail systematically on tasks requiring abstract reasoning. Existing benchmarks document this failure but cannot say \emph{why} it happens or which cognitive capability is missing. We close this gap by adopting the Relational Match-to-Sample (RMTS) paradigm from comparative and developmental psychology and pairing it with a mechanistic analysis of the model's internals. On a parametrically controlled stimulus set evaluated across frontier API models (GPT, Claude, Gemini) and three open-source families (Qwen3.5, Gemma-4, InternVL3), we identify four levers that shift VLMs toward the relational match---capability tier, model scale, the number of objects per scene, and the absence of per-object stimulus noise---together producing a developmental-like trajectory that mirrors the human \emph{relational shift}. Opening up the model, a per-layer representational similarity analysis and a causal mediation analysis reveal that VLM abstract reasoning is implemented by two competing circuits: an early circuit that organises images by their surface object features, and a late circuit that organises them by their abstract relation. Extending the analysis to ARC-AGI-1, we find that ablating the relational heads identified on RMTS degrades performance more than ablating random heads, indicating that the relational circuit is recruited beyond our controlled stimuli. We hope this mechanism-level view serves as a step toward understanding how abstract reasoning is implemented in VLMs.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07646 [cs.AI] |
| (or arXiv:2610.07646v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07646 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gouki Minegishi [view email]
[v1]
Tue, 6 Oct 2026 02:38:40 UTC (788 KB)
来源:arXiv:cs.AI · arxiv.org