跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Genpei Zhang·· 14 小时前AI 评分44

LLaVA、Qwen2.5-VL、InternVL3 中层可解释性研究:错误中编码答案却与预测脱节

Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null

AI 导读

研究在 LLaVA-1.5-7B、Qwen2.5-VL-7B、InternVL3-8B 三种视觉语言模型上发现,POPE 基准中 68-91% 的错误在中层已编码正确答案,但残差流 patching 在层级别对三种架构均产生 0% 非平凡翻转,Qwen 在 12,600 次 per-head 前向中无一翻转。

正文

View PDF HTML (experimental)

Abstract:Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.
Comments: 13 pages, 4 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.00024 [cs.CV]
  (or arXiv:2610.00024v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.00024

arXiv-issued DOI via DataCite

Submission history

From: Genpei Zhang [view email]
[v1] Tue, 1 Sep 2026 20:15:24 UTC (149 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org