跳到正文
arXiv:cs.AI· Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama·· 4 小时前AI 评分48

Qwen2.5-VL-7B 为何"被篡改却仍答对":VLM 的 train/inference gap 机制解析

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

AI 导读

研究在 Qwen2.5-VL-7B-Instruct 上发现,针对性对抗扰动可将固定目标描述的 teacher-forced 训练损失压至近零,但模型自由生成时仍输出原本正确的描述,作者称之为 train/inference gap。

正文

View PDF HTML (experimental)

Abstract:A targeted adversarial perturbation can drive a vision-language model's (VLM's) teacher-forced training loss for a fixed target caption to near zero, yet the same model, allowed to generate freely, produces the original, correct description with no trace of the target. We call this dissociation the train/inference gap, and give it a precise mechanistic account on Qwen2.5-VL-7B-Instruct using a controlled two-stage PGD attack on 200 held-out COCO images. First, we show that image-level pixel statistics, including a correctly re-implemented, texture-based attackability measure from the CNN robustness literature, have essentially no predictive power over which images are corrupted (best predictor r=-0.050, p=0.484; ridge regression R^2=0.069). Second, using the logit lens, we localise the gap to a single autoregressive step: the rank of the target token, conditioned on the correct first token already being generated, is fixed at exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Third, tracking target-token rank across all 28 LLM decoder layers reveals that the visual encoder corrupts every image's representation by a comparable margin regardless of eventual outcome, but the language model decoder then differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it, past its clean-image baseline, for resistant ones (p<0.001, rank-biserial r=0.579). A linear probe on the merger hidden state separates these two outcomes with AUC=0.858, though we flag a circularity concern in this estimate. Together these results argue that adversarial robustness in autoregressive VLMs is substantially a property of the language decoder's prior, not the visual encoder, with direct implications for where faithfulness evaluations and defenses for deployed VLM systems should be targeted.
Comments: Accepted at the VLM4RWD Workshop (Grounded and Faithful Vision-Language Models for Real-World Deployment), NeurIPS 2026. 8 pages, 2 figures, 3 tables
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03445 [cs.CV]
  (or arXiv:2610.03445v1 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2610.03445

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Arun Josephraj Arokiaraj [view email]
[v1] Fri, 2 Oct 2026 15:27:28 UTC (729 KB)

来源:arXiv:cs.AI · arxiv.org