arXiv:cs.LG· Pietro Moriello, Pietro Buzzega, Angelo Porrello, Simone Calderara·· 2 天前AI 评分31
IrekoGPT:将结构化剪枝变为事后可伸缩 LLM
IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs
AI 导读
IrekoGPT 是一种事后方法,可将预训练 LLM 转换为推理时宽度可调的 slimmable 模型。它基于 SliceGPT,保留其投影矩阵不做剪枝,使单一模型暴露不同宽度的嵌套子网络,并通过跨多个压缩比校准每层、用无梯度岭回归修正下游线性层。在 Llama 和 Qwen 模型上,初步结果显示其优于朴素 PCA 剪枝,高压缩率下增益最大。
正文
Abstract:We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression. Across Llama and Qwen models, preliminary results show improvements over naive PCA-based slimming, with the largest gains at high compression. Code is available at this https URL
| Comments: | Accepted at the NeurIPS 2026 Workshop "AXIOM: Foundations of Efficient Deep Learning". 8 pages, 5 figures |
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.00426 [cs.LG] |
| (or arXiv:2610.00426v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00426 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Angelo Porrello [view email]
[v1]
Wed, 30 Sep 2026 16:08:59 UTC (163 KB)
来源:arXiv:cs.LG · arxiv.org