跳到正文
arXiv:cs.AI· Langxi Huang, Pingping Zhang, Lanyun Zhu, Chunyang Jiang, Jiawei Shao, Haocheng Yuan, Peilin Chen·· 6 小时前AI 评分32

ChartBmkAgent:以 harness 管控多智能体,从稀疏错误分类规范构建 Chart QA 基准

ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications

AI 导读

ChartBmkAgent 能从稀疏的错误分类规范出发,由 harness 管控多个专用智能体,自动构建完整的 Chart QA 诊断样本。

正文

View PDF HTML (experimental)

Abstract:Multimodal large language models (MLLMs) advance rapidly, while conventional benchmark development lags behind, delaying investigation of newly observed capability gaps. Such investigation requires an expressive task format and an on-demand construction process: information-rich charts make chart question answering (Chart QA) suitable for probing coupled perception and reasoning. Automated Chart QA construction is intended to shorten the benchmark-development cycle by turning identified gaps into targeted samples on demand. Current methods, however, commonly separate target guidance from scratch generation: target-guided systems often require prepared data, charts, or templates, while scratch-generation systems primarily ensure artifact validity, without explicitly controlling whether newly synthesized requirements and content remain aligned with an externally specified diagnostic target. We introduce ChartBmkAgent, which turns an identified capability gap into targeted diagnostic evidence by constructing complete Chart QA samples from sparse error-taxonomy specifications. Throughout construction, a central harness governs specialized agents, requires stage-specific evidence of alignment with the original error category, and records the basis for each acceptance decision. On 300 taxonomy-wide samples, MLLM accuracies ranged from 32.7% to 84.3% with distinct category profiles, showing that generated samples reveal capability differences. Across three source-model comparisons, targeted follow-ups scored 50.0% versus 82.2% on matched controls ($p=8.96\times10^{-6}$); all six cross-model comparisons had the same direction, demonstrating targeted validation and diagnostic-data generation. Multiple evaluator models assessed whether each sample tested its specified error category; 86.4% met this criterion, providing empirical evidence of target preservation.
Comments: 9 pages, 2 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08106 [cs.AI]
  (or arXiv:2610.08106v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08106

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Langxi Huang [view email]
[v1] Tue, 6 Oct 2026 10:31:28 UTC (357 KB)

来源:arXiv:cs.AI · arxiv.org