arXiv:cs.LG(机器学习,全量分类)· Genpei Zhang·· 14 小时前AI 评分44
LLaVA、Qwen2.5-VL、InternVL3 中层可解释性研究:错误中编码答案却与预测脱节
Encoded but Disconnected: Decomposing Vision-Language Model Failures under a Patching Null
AI 导读
研究在 LLaVA-1.5-7B、Qwen2.5-VL-7B、InternVL3-8B 三种视觉语言模型上发现,POPE 基准中 68-91% 的错误在中层已编码正确答案,但残差流 patching 在层级别对三种架构均产生 0% 非平凡翻转,Qwen 在 12,600 次 per-head 前向中无一翻转。
正文
Abstract:Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers encode the ground-truth answer in 68-91% of errors, yet this signal is not causally active for the final prediction: residual-stream patching yields 0% non-trivial flip at the layer level on all three architectures, and on two of three at the per-head level (Qwen: 0/12,600 patched forwards). The lone exception, InternVL3 layer-20 head-2, is a non-vocab, self-attending head whose effect is localized to that specific head (p < 1e-4). Despite the null, the errors separate operationally into three failure modes -- Perception Failure, Encoded-but-Disconnected, Prior-Override -- learnable above 60% on all three architectures, and the architecture's prior direction predicts which of two interventions elicits a category-specific response. We report these mitigation effects under oracle labels as evidence the categories are mechanistically real, not as a deployable method.
| Comments: | 13 pages, 4 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00024 [cs.CV] |
| (or arXiv:2610.00024v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00024 arXiv-issued DOI via DataCite |
Submission history
From: Genpei Zhang [view email]
[v1]
Tue, 1 Sep 2026 20:15:24 UTC (149 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org