跳到正文
arXiv:cs.AI· Md. Shaown Miah, S. M. Taiabul Haque, Syed Ishtiaque Ahmed, Sayeed Shafayet Chowdhury·· 4 小时前AI 评分59

arXiv 论文提出 TC-LIA 方法检测视觉语言模型的 Mirage 幻觉回答

Detect Before You Leap: Mirage Detection in Vision-Language Models

AI 导读

arXiv 论文提出发布前 Mirage 检测方法 TC-LIA,判断 VLM 回答是否缺乏相关视觉证据而应扣留。该方法在冻结的 CLIP ViT-H/14 编码器各层追踪问题-图像对齐,部署时无需训练。

正文

View PDF HTML (experimental)

Abstract:Vision-language models (VLMs) can produce confident answers without relevant visual evidence, a failure mode known as mirage (Asadi et al., 2026). We study pre-release mirage detection: deciding whether a VLM answer should be released or withheld. Our model-agnostic method, Text-Conditioned Layer-wise Internal Alignment (TC-LIA), tracks question-image alignment across the layers of a frozen CLIP ViT-H/14 encoder, summarizing patch-text alignment by final similarity, late-layer top-k alignment, early-to-late gain, and slope. TC-LIA is training-free at deployment with fixed projections and scoring weights, without any label-specific training, and delivers strong detection independently. Additionally, when combined with blank/noise detection, domain routing, and VLM self-assessment, it forms an ensemble whose supervised training improves performance but is an optional add-on. On 19,004 samples spanning 10 VQA domains, 14 state-of-the-art VLMs exhibit 57.3-75.0% base mirage rates. Our proposed TC-LIA alone cuts this to 7.5% with 83.5% Related/Unrelated/Blank-Noise classification accuracy, and the ensemble reaches 84.5-88.4% accuracy with 5.9-7.2% mirage rates (best joint result: 88.4% accuracy, 6.4% mirage rate). Notably, an ensemble trained on a single backbone transfers well to unseen backbones, with the best-transferring source staying within 0.7% accuracy points of per-backbone training across 13 held-out VLMs.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.00435 [cs.CV]
  (or arXiv:2606.00435v5 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2606.00435

arXiv-issued DOI via DataCite

Submission history

From: Md. Shaown Miah [view email]
[v1] Fri, 29 May 2026 23:51:35 UTC (11,848 KB)
[v2] Mon, 15 Jun 2026 15:09:01 UTC (13,483 KB)
[v3] Mon, 27 Jul 2026 23:31:43 UTC (16,234 KB)
[v4] Wed, 16 Sep 2026 06:47:00 UTC (16,286 KB)
[v5] Thu, 1 Oct 2026 21:23:02 UTC (16,701 KB)

来源:arXiv:cs.AI · arxiv.org