arXiv:cs.AI· Masha Fedzechkina, Eleonora Gualdoni, Rita Ramos, Sinead Williamson·· 4 小时前AI 评分51
研究比较视觉语言模型不同表征层级的信息泄露风险
What do your logits know?
AI 导读
arXiv 论文(arXiv:2604.09885)以视觉语言模型为测试对象,首次系统比较了残差流信息在压缩过程中不同表征层级的信息留存情况,比较对象包括 tuned lens 得到的低维投影和最终 top-k logits 两个自然瓶颈。
正文
Abstract:Recent work has shown that probing model internals can reveal a wealth of information not apparent from the model generations. This poses a risk of unintentional or malicious information leakage, where model users are able to learn information that the model owner assumed was inaccessible. Using vision-language models as a testbed, we present the first systematic comparison of information retained at different representational levels as it is compressed from the rich information encoded in the residual stream through two natural bottlenecks: low-dimensional projections of the residual stream obtained using tuned lens, and the final top-k logits most likely to impact model's answer. We show that even easily accessible bottlenecks defined by the model's top logit values can leak task-irrelevant information present in an image-based query, in some cases revealing as much information as direct projections of the full residual stream.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2604.09885 [cs.AI] |
| (or arXiv:2604.09885v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2604.09885 arXiv-issued DOI via DataCite |
Submission history
From: Masha Fedzechkina [view email]
[v1]
Fri, 10 Apr 2026 20:24:38 UTC (20,784 KB)
[v2]
Fri, 2 Oct 2026 00:50:16 UTC (21,790 KB)
来源:arXiv:cs.AI · arxiv.org