跳到正文
arXiv:cs.LG· Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara·· 2 天前AI 评分31

IrekoGPT:将结构化剪枝变为事后可伸缩 LLM

IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs

AI 导读

IrekoGPT 是一种事后方法,可将预训练 LLM 转换为推理时宽度可调的 slimmable 模型。它基于 SliceGPT,保留其投影矩阵不做剪枝,使单一模型暴露不同宽度的嵌套子网络,并通过跨多个压缩比校准每层、用无梯度岭回归修正下游线性层。在 Llama 和 Qwen 模型上,初步结果显示其优于朴素 PCA 剪枝,高压缩率下增益最大。

正文

View PDF HTML (experimental)

Abstract:We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at this https URL
Comments: Accepted at the NeurIPS 2026 Workshop "AXIOM: Foundations of Efficient Deep Learning". 8 pages, 5 figures
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2610.00426 [cs.LG]
  (or arXiv:2610.00426v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00426

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Angelo Porrello [view email]
[v1] Wed, 30 Sep 2026 16:08:59 UTC (163 KB)

来源:arXiv:cs.LG · arxiv.org