跳到正文
arXiv:cs.CL· Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang·· 3 小时前AI 评分33

面向高效 LLM/VLM 的激活感知帕累托引导低秩压缩

Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM

AI 导读

研究提出低秩压缩框架 PGSVD,通过逐层激活压缩误差上界网络损失变化,并将低秩压缩建模为双目标优化,证明单一统一容差可产生代理帕累托最优的异构秩。该零样本流程结合帕累托引导秩选择与交替最小二乘实现,应用于 LLM 和 VLM 后,在相同压缩率下取得更高精度并实现推理加速。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLM) and vision-language models (VLM) have achieved state-of-the-art performance, but they impose significant memory and computing challenges in deployment. We present a novel low-rank compression framework to address this challenge. First, we upper bound the change of network loss via layer-wise activation-based compression errors, filling a theoretical gap in the literature. We then formulate low-rank model compression as a bi-objective optimization and prove that a single uniform tolerance yields surrogate Pareto-optimal heterogeneous ranks. Based on our theoretical insights, we propose Pareto-Guided Singular Value Decomposition (PGSVD), a zero-shot pipeline that improves activation-aware compression via Pareto-guided rank selection and alternating least-squares implementation. We apply PGSVD to both LLM and VLM, showing better accuracy at the same compression levels and inference speedup.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2510.05544 [cs.CL]
  (or arXiv:2510.05544v3 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2510.05544

arXiv-issued DOI via DataCite

Submission history

From: Ryan Solgi [view email]
[v1] Tue, 7 Oct 2025 03:07:47 UTC (407 KB)
[v2] Wed, 3 Jun 2026 18:20:52 UTC (408 KB)
[v3] Tue, 6 Oct 2026 20:51:02 UTC (405 KB)

来源:arXiv:cs.CL · arxiv.org