跳到正文
arXiv:cs.LG· Mahshad Lotfinia, Sebastian Ziegelmayer, Lisa Adams, Tri-Thien Nguyen, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh·· 2 天前AI 评分63

胸片视觉语言模型未必真的在看图:arXiv 论文提出干预式审计方法

Vision-language models for chest radiography do not always need the image

AI 导读

arXiv 论文对八个开源胸片问答视觉语言模型做干预式审计,通过换图、遮挡、去图等方式检验模型是否真正使用影像。在 MIMIC-CXR 的 2,548 个是非题上,一个多模态模型不看图也答 Yes,仅接收问题文本的医学模型在合并题集上得 55.3%,高于两个多模态系统;在图像必要的场景中,最佳多模态系统比该模型高 10.4% 平衡准确率。

正文

View PDF HTML (experimental)

Abstract:Vision-language models that answer questions about chest radiographs are evaluated by their accuracy on labels derived from radiology reports. High benchmark accuracy is often interpreted as evidence that the model uses the image. A model that answers from the finding named in the question can score as well as a model that uses the radiograph. Keeping the question fixed, we audit eight open-weight systems by swapping in another patient's radiograph with the same or the opposite label, occluding the radiologist-marked region or an equal region elsewhere, and removing the radiograph or replacing it with noise or a photograph. On 2,548 yes-or-no questions from MIMIC-CXR, one multimodal model answers Yes regardless of the image, another multimodal model changes its answers without following the label, and four systems use the image but keep about half of their correct answers when the radiograph is swapped for an opposite-label radiograph. A medical model that receives only the question text scores 55.3% on the pooled questions, higher than two multimodal systems. It scores 91.8% where every finding is present, and answering Yes to every question scores 100% there. Where the image is necessary, the best multimodal system exceeds this model by 10.4% in balanced accuracy. The categories are unchanged on CheXpert. Confidence is not higher when a correct answer depends on the marked region. In a reader study with three radiologists, the two radiologists who read a balanced set of 200 cases score 86.0% and 82.0%, and the systems score 50.0% to 73.0%. Accuracy does not establish image use, but an intervention on the image can test it.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2606.17710 [cs.CV]
  (or arXiv:2606.17710v3 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2606.17710

arXiv-issued DOI via DataCite

Submission history

From: Soroosh Tayebi Arasteh [view email]
[v1] Tue, 16 Jun 2026 09:22:10 UTC (599 KB)
[v2] Fri, 19 Jun 2026 19:06:31 UTC (599 KB)
[v3] Thu, 1 Oct 2026 11:48:16 UTC (1,065 KB)

来源:arXiv:cs.LG · arxiv.org