跳到正文
arXiv:cs.CL· Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Bangrui Xu, Yuhan Wang, Bin Wang, Conghui He·· 3 小时前

CiteVQA:面向可信文档智能的证据归因基准测试

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence

AI 导读

多模态大模型文档理解评测 CiteVQA 发布,要求模型在给出答案的同时返回元素级 bounding-box 引用,并联合评估两者。该基准包含 711 份 PDF 上的 1,897 道问题,覆盖七个领域和两种语言,平均每份文档 40.6 页,采用严格归因准确率(SAA)指标——答案与引用区域均正确才计分。

正文

Authors:Dongsheng Ma, Jiayu Li, Zhengren Wang, Yijie Wang, Jiahao Kong, Weijun Zeng, Jutao Xiao, Jie Yang, Bangrui Xu, Yuhan Wang, Bin Wang, Conghui He

View PDF HTML (experimental)

Abstract:Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only approach masks a critical failure mode: a model can land on the correct answer while grounding it in the wrong passage---a critical risk in high-stakes domains like law, finance, and medicine, where every conclusion must be traceable to a specific source region. To address this, we introduce CiteVQA, a benchmark that requires models to return \textit{element-level} bounding-box citations alongside each answer, evaluating both jointly. CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, averaging 40.6 pages per document. To ensure fidelity and scalability, the ground-truth citations are generated by an automated pipeline---which identifies crucial evidence via masking ablation and enforces multi-stage quality control. At the core of our evaluation is Strict Attributed Accuracy (SAA), which credits a prediction only when the answer and the cited region are both correct. Auditing 20 MLLMs reveals a pervasive Attribution Hallucination: models frequently produce the right answer while citing the wrong region. The strongest system (Gemini-3.1-Pro-Preview) achieves an SAA of only 76.0, and the strongest open-source MLLM reaches just 22.5. Ultimately, towards trustworthy document intelligence, CiteVQA exposes a reliability gap that answer-only evaluations overlook, providing the instrumentation needed to close it. Our repository is available at this https URL.
Subjects: Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2605.12882 [cs.CL]
  (or arXiv:2605.12882v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.12882

arXiv-issued DOI via DataCite

Submission history

From: Dongsheng Ma [view email]
[v1] Wed, 13 May 2026 01:54:42 UTC (4,251 KB)
[v2] Thu, 8 Oct 2026 06:54:20 UTC (4,042 KB)

来源:arXiv:cs.CL · arxiv.org