跳到正文
arXiv:cs.CL· Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin, Akiko Aizawa·· 4 小时前AI 评分41

多模态 LLM 为何在图表证据上表现更差:编码了但没被路由

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

AI 导读

研究通过逐层线性探针与注意力分析发现,多模态 LLM 在科学论断验证中对图表的处理并非编码失败,而是图表信息虽被编码进中间表示,却未能到达预测位置,这一断连在表格证据中不存在。该现象在三个开源权重 VLM 及所有测试条件下均成立,且注意力分析显示不同模型家族存在两种架构上不同的断连形式。

正文

View PDF HTML (experimental)

Abstract:Multimodal LLMs are increasingly used to assist scientific peer review, where a core requirement is verifying whether claims in a paper are supported by its evidence. Prior work has shown that models perform substantially better at this task when the evidence is a table than when it is a chart of the same underlying data. This raises the question of whether models fail to extract information from charts, or do they extract it but fail to use it when forming their prediction? We study this question through layer-wise linear probing and attention analysis on three open-weight VLMs over table and chart evidence, representing the same underlying data. We find consistent evidence for the latter. Chart information is encoded in the models' intermediate representations but does not reach the prediction position, a gap that is absent for tables and holds across all conditions tested. Attention analysis further reveals that this disconnect takes two architecturally distinct forms across model families. These findings point toward reframing the table-chart gap as a failure of how encoded visual information is used at prediction time, rather than a failure of encoding itself.
Comments: Accepted to AACL-IJCNLP 2026 Findings
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2606.01679 [cs.CL]
  (or arXiv:2606.01679v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.01679

arXiv-issued DOI via DataCite

Submission history

From: Sunisth Kumar [view email]
[v1] Mon, 1 Jun 2026 04:39:13 UTC (264 KB)
[v2] Fri, 2 Oct 2026 05:39:03 UTC (265 KB)

来源:arXiv:cs.CL · arxiv.org