跳到正文
arXiv:cs.LG· Mehdi Makni, Xiang Meng, Rahul Mazumder·· 3 小时前

3BASiL-TM:面向 LLM 稀疏加低秩压缩的算法框架

3BASiL: An Algorithmic Framework for Sparse plus Low-Rank Compression of LLMs

AI 导读

3BASiL-TM 是一种一次性的训练后 LLM 稀疏加低秩(S+LR)分解方法,在 (2:4 Sparse + 64 LR) 配置下将相对稠密 LLaMA-8B 的 WikiText2 困惑度差距缩小超过 30%,并在 A100 GPU 上实现超过 2.5 倍的压缩运行速度。

正文

View PDF HTML (experimental)

Abstract:Sparse plus Low-Rank $(\mathbf{S} + \mathbf{LR})$ decomposition of Large Language Models (LLMs) has emerged as a promising direction in model compression, aiming to decompose pre-trained model weights into a sum of sparse and low-rank matrices $(\mathbf{W} \approx \mathbf{S} + \mathbf{LR})$. Despite recent progress, existing methods often suffer from substantial performance degradation compared to dense models. In this work, we introduce 3BASiL-TM, an efficient one-shot post-training method for $(\mathbf{S} + \mathbf{LR})$ decomposition of LLMs that addresses this gap. Our approach first introduces a novel 3-Block Alternating Direction Method of Multipliers (ADMM) method, termed 3BASiL, to minimize the layer-wise reconstruction error with convergence guarantees. We then design an efficient transformer-matching (TM) refinement step that jointly optimizes the sparse and low-rank components across transformer layers. This step minimizes a novel memory-efficient loss that aligns outputs at the transformer level. Notably, the TM procedure is universal as it can enhance any $(\mathbf{S} + \mathbf{LR})$ decomposition, including pure sparsity. Our numerical experiments show that 3BASiL-TM reduces the WikiText2 perplexity gap relative to dense LLaMA-8B model by over 30% under a (2:4 Sparse + 64 LR) configuration, compared to prior methods. Moreover, our method achieves over 2.5x faster compression runtime on an A100 GPU compared to SOTA $(\mathbf{S} + \mathbf{LR})$ method. Our code is available at this https URL.
Comments: The Thirty-ninth Annual Conference on Neural Information Processing Systems
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2603.01376 [cs.LG]
  (or arXiv:2603.01376v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2603.01376

arXiv-issued DOI via DataCite

Submission history

From: Mehdi Makni [view email]
[v1] Mon, 2 Mar 2026 02:16:46 UTC (1,261 KB)
[v2] Thu, 8 Oct 2026 00:16:32 UTC (1,562 KB)

来源:arXiv:cs.LG · arxiv.org