跳到正文
arXiv:cs.CL· Ruochi Li, Jianzhe Lin, Haoxuan Zhang, Haihua Chen, Junhua Ding, Edward Gehringer, Yang Zhang·· 3 小时前AI 评分41

TopoGraphRAG-Bench:评估多模态 GraphRAG 在版面证据推理上的表现

TopoGraphRAG-Bench: Evaluating Multimodal GraphRAG on Layout-Grounded Evidence Reasoning

AI 导读

研究者推出 TopoGraphRAG-Bench,一个面向多模态 GraphRAG 的版面证据推理基准,包含 201 篇长文档上的 2,024 道题,按单跳检索、桥接链推理和多源综合三种拓扑构建。评测显示多模态 GraphRAG 整体表现最强,但在视觉-文本证据对齐或多单元组合不完整时仍会失败;纯文本 GraphRAG 在关键依赖落在图表时表现吃力,页面级视觉检索则缺乏拓扑恢复所需的细粒度结构。

正文

View PDF HTML (experimental)

Abstract:Real-world documents distribute evidence across text, tables, figures, and captions within complex page layouts. Answering complex questions over such documents therefore requires more than retrieving relevant passages: systems must recover the evidence topology that connects heterogeneous evidence units. Existing GraphRAG evaluations remain largely text-centered, while multimodal document RAG benchmarks assess cross-modal retrieval and generation without directly evaluating recovery of the intended evidence topology. We introduce TOPOGRAPHRAG-BENCH, a layout-grounded benchmark for multimodal evidence reasoning in GraphRAG, comprising 2,024 questions over 201 long, visually rich documents. Questions are constructed bottom-up from text, figure, and table evidence units under three controlled topologies: single-hop retrieval, bridge-chain reasoning, and multi-source synthesis. To ensure that questions preserve their intended structure, we apply counterfactual validation for shortcut resistance, modality necessity, and evidence necessity. We evaluate text-only GraphRAG, page-level visual retrieval, and multimodal GraphRAG systems using retrieval, generation, and topology-aware reasoning metrics. Multimodal GraphRAG systems achieve the strongest overall performance, but still fail when visual-textual evidence alignment or multi-unit composition is incomplete. Text-only GraphRAG struggles when key dependencies are grounded in figures or tables, while page-level visual retrieval lacks the fine-grained structure needed for topology recovery. These findings motivate GraphRAG systems that move beyond text-derived entity relation graphs to explicitly model document layouts, cross-modal evidence alignment, and the reasoning roles of evidence units. Code and data are available at this https URL.
Comments: Accepted at the 40th Conference on Neural Information Processing Systems (NeurIPS 2026)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.09360 [cs.AI]
  (or arXiv:2610.09360v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.09360

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruochi Li [view email]
[v1] Wed, 7 Oct 2026 03:15:31 UTC (1,526 KB)

来源:arXiv:cs.CL · arxiv.org