arXiv:cs.LG· Ambuj Mehrish, Sebastiano Vascon·· 6 小时前AI 评分34
FlowReader:面向长多模态文档证据组装的最小成本流路由
Min-Cost Flow Routing for Evidence Assembly in Long Multimodal Documents
AI 导读
FlowReader 将长多模态文档的证据选择建模为多模态内容图上的单一最小成本流问题,通过谱分解按谱能量分配证据预算,无需语言模型规划调用即可强制覆盖各查询相关方面。
正文
Abstract:Answering questions about long multimodal documents requires distributing a fixed evidence budget across relevant facets in text, tables, figures, and slides while avoiding near-duplicates. We present \flowreader, which formulates evidence selection as a single minimum-cost flow problem with capacity limits over a multimodal content graph. Spectral decomposition identifies latent aspects of query-relevant content and allocates the budget among them in proportion to their spectral energy. These capacity limits enforce aspect coverage during routing without requiring a language-model planning call. Query-conditioned costs prioritize chains of relevant, mutually consistent evidence. Decomposing the optimal flow produces short evidence chains, which a vision-language model reads in parallel and a reasoner reconciles. On VisDoMBench with Qwen3-VL-32B, \flowreader\ achieves the highest macro accuracy ($68.9$), surpassing the strongest prior system by $2.7$ points, leading on three of five subsets and attaining the highest worst-subset accuracy. It uses a measured $17.5$ content nodes per query and maintains its lead at $12.9$. Ablation studies with a fixed graph, scorer, reader, and judge show that cost design drives accuracy, capacity limits preserve it while using about three-quarters of the reader tokens required by shortest-path routing without these limits on the same network, and spectral aspects align with LLM-generated sub-questions without a planning call.
| Subjects: | Information Retrieval (cs.IR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.07235 [cs.IR] |
| (or arXiv:2606.07235v3 [cs.IR] for this version) | |
| https://doi.org/10.48550/arXiv.2606.07235 arXiv-issued DOI via DataCite |
Submission history
From: Ambuj Mehrish [view email]
[v1]
Fri, 5 Jun 2026 13:05:00 UTC (2,919 KB)
[v2]
Mon, 8 Jun 2026 12:31:09 UTC (2,919 KB)
[v3]
Fri, 2 Oct 2026 08:18:22 UTC (2,126 KB)
来源:arXiv:cs.LG · arxiv.org