arXiv:cs.LG(机器学习,全量分类)· Ivan Ilin, Peter Richt\'arik·· 15 小时前AI 评分36
QK-Wanda:耦合 Query 与 Key 的非结构化剪枝方法
QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
AI 导读
QK-Wanda 通过未掩码的 pre-RoPE 重建目标,按单独删除代价对 query 和 key 权重打分,并引入对侧投影信息,使两个投影共享剪枝预算。该方法无需梯度或权重更新,在 A100 上完整剪枝仅比 Wanda 慢 1.3%,H200 上慢 3.1%。
正文
Abstract:Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Llama 2, Llama 3, and Qwen2.5, spanning 0.5B-72B parameters. Relative to Wanda, QK-Wanda reduces QK reconstruction error by an average of 60% at 50% sparsity and 45% at 80%. Downstream gains depend on the model. At 80% sparsity on Llama 2 70B, WikiText-2 and C4 perplexity decrease by 20.3% and 13.5%, while mean zero-shot accuracy rises by 5.94 percentage points. Qwen2.5-72B also improves, but Llama-3.1-70B has substantially higher perplexity despite lower reconstruction error. These results show both the promise of coupled pruning criteria and the limits of local reconstruction as a predictor of model quality.
| Comments: | 81 pages, including appendices |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.01554 [cs.LG] |
| (or arXiv:2610.01554v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01554 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ivan Ilin [view email]
[v1]
Thu, 1 Oct 2026 12:22:44 UTC (688 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org