arXiv:cs.LG· Subba Reddy Oota, Anant Khandelwal, Khushbu Pahwa, Satya Sai Srinath Namburi, Tanmoy Chakraborty, Bapi S. Raju, Manish Gupta·· 6 小时前AI 评分34
视觉语言模型与动作模型在自然游戏中的推理与动作表征脑对齐研究
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
AI 导读
研究利用fMRI记录被试玩Atari风格游戏时的脑活动,对比视觉语言模型(VLMs)与大动作模型(LAMs)的脑对齐表现。两类模型在体素级编码性能上均显著优于强化学习基线,即使在特征维度匹配下仍保持优势,且提示词驱动的提升在额顶叶与运动规划区域约为2-2.5倍。
正文
Abstract:Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model types, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape the models' internal representations and their alignment with fMRI brain activity. First, we find that both VLMs and LAMs achieve significantly higher voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, compared to a no-prompt baseline, prompt-driven gains are larger in higher-order frontal-parietal and motor-planning regions than in early visual cortex, roughly 2-2.5$\times$ when averaged over region-of-interest (ROI) groups, although individual regions are heterogeneous. Third, variance partitioning reveals a qualitatively different representational organization. VLM representations are prompt-symmetric (12.4% unique action vs. 9.5% unique reasoning), whereas LAM representations are action-dominant (25.6% unique action vs.-8.2% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results associate action specialization with distinct cortical alignment patterns in multimodal game-state representations, revealing differences hidden by similar prediction accuracy.
| Comments: | 32 pages, 20 figures |
| Subjects: | Neurons and Cognition (q-bio.NC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2605.19352 [q-bio.NC] |
| (or arXiv:2605.19352v2 [q-bio.NC] for this version) | |
| https://doi.org/10.48550/arXiv.2605.19352 arXiv-issued DOI via DataCite |
Submission history
From: Subba Reddy Oota [view email]
[v1]
Tue, 19 May 2026 04:40:14 UTC (2,789 KB)
[v2]
Wed, 7 Oct 2026 08:01:15 UTC (3,105 KB)
来源:arXiv:cs.LG · arxiv.org