arXiv:cs.AI· Faezeh Dehghan Tarzjani, Mevan Wijewardena, Alexander Romanus, Sampad Mohanty·· 6 小时前AI 评分32
从局部证据到安全判定:视觉语言模型中的因果追踪
From Local Evidence to Safety Verdicts: Causal Tracing in Vision-Language Models
AI 导读
研究者提出 SSU-Bench 数据集,用单项提示词编辑或图像编辑构建成对的安全与不安全图文组合,并在三个视觉语言模型上交换配对输入的内部状态以测量安全判定变化。
正文
Abstract:A vision-language model may need to combine an image with a prompt to recognize a safety risk that neither reveals alone. Where does this joint safety judgment become accessible inside the model? We introduce SSU-Bench, a dataset of matched safe and unsafe image-text combinations constructed using single-item prompt edits or image edits with annotated intended regions. Using three vision-language models, we transfer internal states between paired inputs and measure the resulting change in the safety verdict. Across models and both types of counterfactual, interventions at the changed input positions are effective in earlier decoder layers, while interventions at the final input token become effective later. Directions estimated from other examples produce similar late-layer effects. A linear readout of the final-token state also predicts the model's own verdict, including incorrect judgments, and cross-model comparisons reveal similarities in the patterns of counterfactual change. These findings identify a recurring transition in where interventions can influence a joint safety verdict and distinguish a readable model decision from a correct safety judgment.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07514 [cs.AI] |
| (or arXiv:2610.07514v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07514 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Faezeh Dehghan Tarzjani [view email]
[v1]
Mon, 5 Oct 2026 23:27:19 UTC (5,792 KB)
来源:arXiv:cs.AI · arxiv.org