跳到正文
arXiv:cs.LG· Tianxiao Cao, Jiahe Shao, Yuning Qiu, Kyohei Atarashi, Hisashi Kashima, Qibin Zhao·· 4 小时前AI 评分42

SLBF:面向 MoE 大语言模型的无数据低秩基分解压缩方法

Shared Low-rank Basis Factorization for Data-free Mixture-of-Experts Compression

AI 导读

Shared Low-rank Basis Factorization(SLBF)是一种无需数据的 MoE 权重重构压缩方法,通过专家间共享 rank-k 基实现更丰富的跨专家共享与更低重构误差,并用事后 gauge fixing 无损去除冗余参数。

正文

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) large language models decouple capacity from compute through sparse routing, but their large parameter count creates storage and serving challenges. We analyze three MoE compression families: expert pruning, expert merging, and weight reconstruction, and derive structural error bounds showing that pruning and merging can incur non-vanishing errors tied to routing and expert heterogeneity. In contrast, weight reconstruction avoids these structural costs by preserving expert structure and routing. Motivated by the analysis, we propose Shared Low-rank Basis Factorization (SLBF), a data-free weight reconstruction method that uses rank-$k$ bases shared among experts, enabling richer cross-expert sharing, faster convergence, and lower reconstruction error. A post-hoc gauge fixing removes redundant parameters at no representational cost. Across five MoE architectures spanning 16B to 122B parameters, SLBF consistently outperforms methods from all three compression families.
Comments: Accepted to Findings of EMNLP 2026
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.09342 [cs.LG]
  (or arXiv:2610.09342v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09342

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Tianxiao Cao [view email]
[v1] Wed, 7 Oct 2026 03:01:19 UTC (139 KB)

来源:arXiv:cs.LG · arxiv.org