arXiv:cs.LG· Tianxiao Cao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, Qibin Zhao·· 4 小时前AI 评分42
SLBF:面向 MoE 大语言模型的无数据低秩基分解压缩方法
Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression
AI 导读
Shared Low-rank Basis Factorization(SLBF)是一种无需数据的 MoE 权重重构压缩方法,通过专家间共享 rank-k 基实现更丰富的跨专家共享与更低重构误差,并用事后 gauge fixing 无损去除冗余参数。
正文
Abstract:Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
| Comments: | Accepted to Findings of EMNLP 2026 |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09342 [cs.LG] |
| (or arXiv:2610.09342v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09342 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tianxiao Cao [view email]
[v1]
Wed, 7 Oct 2026 03:01:19 UTC (139 KB)
来源:arXiv:cs.LG · arxiv.org