arXiv:cs.LG· Rahul Chowdhury, Timothy A Rupprecht, Xuan Shen, Shaoyi Huang, Pu Zhao, Yanzhi Wang·· 5 小时前AI 评分44
Patch-to-Prune(P2P):无需训练即可剪枝视觉语言模型中的视觉计算
From Patching to Pruning Visual Computation in Vision Language Models
AI 导读
研究者提出免训练框架 Patch-to-Prune(P2P),通过验证引导的前向与反向层扫描,用固定中性代理激活向量替换部分解码器层的视觉 token 投影输出,在不改变序列长度、token 顺序、位置信息、注意力掩码和残差路径的前提下剪枝计算。
正文
Abstract:Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence length, token order, positional information, attention mask, and residual pathways, thereby pruning computation without removing tokens or modifying the pretrained model weights. We evaluate P2P on four VLMs from the Qwen2.5-VL and LLaVA families across seven multi-modal benchmarks using mutually disjoint calibration, validation, and test partitions. P2P at a 3% tolerance retains around 94% of dense accuracy while reducing FLOPs by 55%. Beyond these efficiency gains, our layer-wise analysis suggests that visual processing in VLMs is non-uniformly distributed across decoder depth: early and late layers often require little token-specific visual computation, whereas intermediate layers appear to perform most task-relevant visual integration, enabling later reasoning to rely largely on visual information already embedded in shared residual and textual representations. This makes P2P both an efficient inference framework and a causal lens into visual information processing in VLMs.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.03389 [cs.CV] |
| (or arXiv:2610.03389v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03389 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Rahul Chowdhury [view email]
[v1]
Fri, 2 Oct 2026 14:37:53 UTC (4,721 KB)
来源:arXiv:cs.LG · arxiv.org