arXiv:cs.LG(机器学习,全量分类)· Xinhe Wu, Yadong Jin·· 5 小时前AI 评分38
ChartDensity-Bench:评测 MLLMs 在视觉密度下的数值数据重建能力
ChartDensity-Bench: Benchmarking MLLMs for Numerical Data Reconstruction under Visual Density
AI 导读
研究者推出 ChartDensity-Bench,用于评测多模态大语言模型(MLLMs)在受控视觉密度下从复合图表中重建结构化数值数据的能力。该基准将同屏图表数量设为 k∈{1,3,6,9},并从结构可靠性、重建完整性、可解析性和数值保真度四个维度评估。对五个近期 MLLMs 的实验显示,数值重建能力随视觉密度上升而普遍下降,且降幅因模型而异。
正文
Abstract:Multimodal large language models (MLLMs) offer a promising approach for recovering numerical data from scientific charts, but their ability to reconstruct chart data from visually dense figures remains poorly understood. Existing chart understanding benchmarks primarily evaluate question answering or chart-level reasoning and provide limited support for evaluating structured numerical reconstruction from scientific figures. We introduce \textbf{ChartDensity-Bench}, a benchmark for evaluating MLLMs on structured numerical data reconstruction from compound chart figures under controlled visual density. Built from charts paired with source-level ground-truth data, ChartDensity-Bench systematically varies the number of simultaneously presented charts ($k\in{1,3,6,9}$), enabling controlled evaluation of density-induced degradation. We further propose a multi-dimensional evaluation framework covering structural reliability, reconstruction completeness, parseability, and numerical fidelity. Experiments on five recent MLLMs show that numerical reconstruction generally degrades as visual density increases, while the magnitude of degradation varies substantially across models. Chart-level paired comparisons further show that the same source chart can incur higher reconstruction error when embedded in denser visual contexts. These findings highlight visual density as an important and previously underexplored factor in MLLM chart data reconstruction and provide a systematic benchmark for evaluating model robustness in this setting.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38781 [cs.LG] |
| (or arXiv:2609.38781v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38781 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinhe Wu [view email]
[v1]
Wed, 30 Sep 2026 02:13:14 UTC (258 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org