arXiv:cs.AI· Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin·· 4 小时前AI 评分58
CRAFT:定位医疗视觉语言模型中因果责任与失败轨迹的机制分析
CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models
AI 导读
研究提出 CRAFT 框架,将医疗视觉语言模型的两种安全失败模式定位到最小因果注意力头集合:文本覆盖视觉的仲裁失败由中深层宽带仲裁头介导,证据不足仍作答的刹车失败由中后层窄带刹车头介导。
正文
Abstract:As vision language models are increasingly deployed in clinical diagnosis, under standing how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitra tion failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores ap propriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and inter ventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at this https URL.
| Comments: | NeurIPS 2026 Spotlight, Medical VLM Failure Analysis |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38810 [cs.CV] |
| (or arXiv:2609.38810v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38810 arXiv-issued DOI via DataCite |
Submission history
From: Jiaqi Zeng [view email]
[v1]
Wed, 30 Sep 2026 02:41:33 UTC (4,678 KB)
[v2]
Fri, 2 Oct 2026 15:11:54 UTC (4,678 KB)
来源:arXiv:cs.AI · arxiv.org