arXiv:cs.LG· Shuailong Wang, Xinyu Lyu, Shengming Yuan, Jingkuan Song, Heng Tao Shen, Lianli Gao·· 4 小时前AI 评分45
理解与缓解 Token-Pruning 在 VLM 中引发的安全漏洞
Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMs
AI 导读
研究首次系统评估 Token-Pruning 对 VLM 安全性的影响,发现多数剪枝策略随比例升高显著降低安全性,而 Query-based Compression 在极端剪枝(最高 99.8%)下反而提升安全性。
正文
Abstract:Token-Pruning accelerates Vision-Language Models by removing redundant visual tokens, yet its safety implications remain underexplored. In this work, we present the first comprehensive safety evaluation of Token-Pruning mechanisms and find that: most pruning strategies significantly degrade safety as pruning ratios increase, whereas Query-based Compression shows the opposite, with extreme pruning (up to 99.8%), unexpectedly improves model safety. This sharp contrast prompts a key question: How do different Token-Pruning strategies reshape model safety behavior, and is it possible to enhance safety without sacrificing acceleration? To answer this, we identify an unrecognized mechanism, termed Pruning-Induced Malicious Amplification, where removal of background tokens triggers a side effect: forcing the model's attention to collapse onto a few retained malicious anchors within the foreground, inadvertently amplifying their toxic semantics under jailbreak. To address that, we propose an inference-time and plug-and-play Safety-Aware Pruning (SAP) mechanism that counteracts such dominance via three steps: (1) identifying malicious anchors, (2) restoring pruned benign tokens, and (3) reallocating excessive attention from malicious anchors to benign tokens. Extensive experiments across three safety and four utility benchmarks demonstrate that SAP mitigates pruning-induced vulnerabilities, i.e., reducing ASR by up to 62%, without compromising efficiency or utility.
| Comments: | Accepted at ICML 2026. 16 pages |
| Subjects: | Cryptography and Security (cs.CR); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09703 [cs.CR] |
| (or arXiv:2610.09703v1 [cs.CR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09703 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:128879-128894, 2026 |
Submission history
From: Shuailong Wang [view email]
[v1]
Wed, 7 Oct 2026 09:01:09 UTC (10,177 KB)
来源:arXiv:cs.LG · arxiv.org