跳到正文
arXiv:cs.AI· Bach Nguyen, Zhaonan Li, Mau Son Nguyen, Sanika Chavan, Nilay Kumar, Hong Anh Nguyen, Khoa Vo, Ben Zhou·· 6 小时前AI 评分40

VisualNoiseQA:LLM 如何在噪声证据下推理?一个主动视觉推理基准

How Well Do LLMs Reason with Noisy Evidence? An Active Visual Reasoning Benchmark

AI 导读

研究者提出 VisualNoiseQA,一个面向噪声视觉反馈下主动推理的基准:纯文本 LLM 需通过迭代查询作为随机视觉传感器的现成 VLM 来解答 VQA 问题,每次查询用自一致性给出经验不确定性信号。基准覆盖感知、图表理解和知识密集型推理共 1,000 道题,仅保留传感器回答不一致但人类可解的问题。

正文

View PDF HTML (experimental)

Abstract:Real-world reasoning rarely reduces to static question answering: agents must actively gather information from tools and sensors that are often noisy and unreliable. Yet most existing active reasoning benchmarks assume that environmental feedback is trustworthy, or introduce noise without exposing an explicit, calibrated uncertainty signal, leaving open how LLMs should reason when the evidence itself is uncertain. We introduce VisualNoiseQA, a novel benchmark for active reasoning under noisy visual feedback. A text-only LLM must solve VQA problems by iteratively querying a fixed, off-the-shelf VLM treated as a stochastic visual sensor. For each query, we draw multiple samples and expose an empirical uncertainty signal via self-consistency, enabling the reasoner to probe from different angles and decide what to ask next and when to stop. Our construction is automatic and scalable: starting from diverse VQA sources and two noisy VLMs, we retain only questions where the sensor is inconsistent yet human-solvable. We evaluate multiple LLM reasoners on 1,000 instances spanning perception, chart understanding, and knowledge-intensive reasoning. VisualNoiseQA thus provides a controlled playground to study how different LLMs exploit uncertainty signals for robust reasoning.
Comments: 27 pages, 9 figures, 11 tables
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.07751 [cs.AI]
  (or arXiv:2610.07751v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07751

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bach Nguyen [view email]
[v1] Tue, 6 Oct 2026 04:49:11 UTC (3,576 KB)

来源:arXiv:cs.AI · arxiv.org