arXiv:cs.CL· Yupeng Xie, Zhenyang Wang, Jiayi Zhu, Yinghao Tang, Zhouan Shen, Yiyu Chen, Yuyu Luo·· 3 小时前
DataVista:首个数据视频理解基准,诊断多模态 LLM 的图表叙事能力
DataVista: Diagnosing Multimodal LLMs on Data Video Understanding
AI 导读
研究者发布 DataVista,首个面向数据视频理解的基准,含 961 个真实数据视频与 6,775 道评测题,按数据感知、时间推理、叙事理解三级能力框架和 10 种题型组织。对 19 个主流 MLLM 的评测显示,表现最佳的 Gemini-3.1-Pro 总体准确率仅 70.0%,远低于人类专家,且在因果推理与叙事结构上最弱。增加帧数和字幕主要提升数据感知与时间推理,对叙事理解增益有限。
正文
Abstract:Data video is a media form that integrates data visualization with video narrative, widely adopted in news reporting and business analysis. Compared with general video understanding, data video understanding places greater emphasis on accurately reading data from animated charts, integrating evidence across charts and time, and understanding how narrative organization and visual design communicate information. Yet existing benchmarks target either general videos or static charts, and data video understanding has not been systematically evaluated. We present DataVista, the first benchmark for data video understanding, containing 961 real-world data videos and 6,775 evaluation questions organized under a three-level progressive capability framework (data perception, temporal reasoning, narrative understanding) with 10 fine-grained question types across five topic domains. Systematic evaluation of 19 mainstream MLLMs shows that the best-performing model, Gemini-3.1-Pro, achieves 70.0% overall accuracy, still far below human expert performance, with models performing worst on Causal Reasoning and Narrative Structure. Increasing frame counts and adding subtitles mainly benefit data perception and temporal reasoning, with limited gains in narrative understanding. Further analysis of model responses identifies typical failure modes in chart reading, evidence judgment, and instruction understanding. The benchmark is available at this https URL.
| Comments: | 46 pages, 22 figures, 14 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11993 [cs.CV] |
| (or arXiv:2610.11993v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11993 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yupeng Xie [view email]
[v1]
Thu, 8 Oct 2026 14:02:35 UTC (15,579 KB)
来源:arXiv:cs.CL · arxiv.org