跳到正文
arXiv:cs.CL· Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal·· 4 小时前AI 评分32

用合作原则评估视觉语言模型的 VQA 表现

Evaluating VQA in Vision Language Models using Cooperative Principles

AI 导读

一项研究用 Grice 准则被违反的问题测试视觉语言模型(VLM)的 VQA 表现,发现 ChatGPT、Claude、Gemini 和 Llava 在问题含冗余、歧义或虚假信息时准确率下降。研究还对比了人类与 VLM 的语用推理差异,以及 VLM 处理人类诱导与 AI 生成违规时的表现差异。人类处理 VLM 诱导违规的认知投入更低,但 VLM 自身在这类情况下准确率更差。

正文

View PDF HTML (experimental)

Abstract:We evaluate the performance of Vision Language Models in Visual Question Answering (VQA) when questions violate Grice's maxims. To do this, we use VLMs to generate question modifiers that add non-essential, ambiguous or false information and show that in the presence of such violations, the VLMs that we evaluate (ChatGPT, Claude, Gemini and Llava) show diminished performance. Further, we empirically show the difference between how humans reason pragmatically compared to VLMs, and the difference in VLM reasoning when it resolves violations that are human-induced compared to those that are AI-generated. Finally, we show that human cognitive effort (measured through time-on-task in an experiment) is lower for resolving VLM-induced violations, but VLMs themselves perform less accurately in such cases.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.02878 [cs.CL]
  (or arXiv:2610.02878v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.02878

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Monika Shah [view email]
[v1] Fri, 2 Oct 2026 06:13:15 UTC (7,102 KB)

来源:arXiv:cs.CL · arxiv.org