arXiv:cs.LG· Shuo Yang, Changbai Li, Linlin Yang, Huobin Tan, Rongyu Chen, Tongfei Chen, Tian Wang, Sheng Xu, Baochang Zhang·· 7 小时前AI 评分30
DIPrune:面向高效多模态语言模型的双重要性任务感知 token 剪枝
DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models
AI 导读
DIPrune 是一种无需训练的多模态大语言模型 token 剪枝框架,通过联合优化层内静态特征显著性与层间动态语义演化来减少计算开销。该方法将剪枝重构为最终任务损失畸变的最小化问题,并推导出 token 级上界作为代理目标,揭示了一个此前被忽视的跨层梯度项。在 LLaVA 和 Qwen-VL 上的实验表明,DIPrune 持续取得 SOTA 结果。
正文
Abstract:Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tractable, token-wise upper bound to serve as a surrogate objective. Specifically, this formulation inherently reveals a previously neglected inter-layer term that accounts for gradients across layers. Accordingly, for the implementation, we propose DIPrune, a rank-based framework that employs a dual importance scoring mechanism to jointly optimize intra-layer static feature saliency and inter-layer dynamic semantic evolution. Extensive experiments on LLaVA and Qwen-VL demonstrate that DIPrune consistently achieves state-of-the-art results.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08341 [cs.CV] |
| (or arXiv:2610.08341v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08341 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuo Yang [view email]
[v1]
Tue, 6 Oct 2026 13:38:27 UTC (3,567 KB)
来源:arXiv:cs.LG · arxiv.org