arXiv:cs.LG(机器学习,全量分类)· Yushuai Sun, Zikun Zhou, Lin Gao, Jun Yu, Wenjie Pei·· 14 小时前AI 评分34
MoRA:通过路由器偏置学习与专家近似实现 MoE 剪枝
MoRA: MoE Pruning via Router Bias Learning and Expert Approximation
AI 导读
研究者提出结构化 MoE 专家剪枝框架 MoRA,为每个专家引入可学习路由器偏置,通过最小化语言建模损失和路由多样性正则项来优化,并加入专家近似机制用剩余专家以仿射变换逼近被剪枝专家的输出。在 Qwen3-30B-A3B、DeepSeek-V2-Lite 和 Moonlight-16B-A3B 上每层移除 25% 和 50% 路由专家,九个零样本基准测试中优于现有剪枝算法,代码将开源。
正文
Abstract:Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. However, existing pruning methods either use expert ranking criteria that are not well aligned with model performance or rely on effective expert subset searching that is computationally expensive. Moreover, these methods typically overlook the routing-behavior redundancy among the retained experts. In this paper, we propose MoE Pruning via Router Bias Learning and Expert Approximation (MoRA), a framework for structured MoE expert pruning. We introduce a learnable router bias for each expert and optimize these biases by minimizing the language-modeling loss and a routing-diversity regularizer. The learned router biases sharpen the routing probability distributions to identify experts critical to model performance while encouraging the selection of experts with diverse routing preferences. In addition, we introduce an expert approximation mechanism as a post-pruning enhancement. It leverages the remaining experts to approximate the outputs of pruned experts by affine transformation, further improving the performance of the pruned model. We evaluate MoRA on Qwen3-30B-A3B, DeepSeek-V2-Lite, and Moonlight-16B-A3B, removing 25\% and 50\% of the routed experts in each MoE layer. Extensive experiments on nine zero-shot benchmarks show that MoRA outperforms state-of-the-art pruning algorithms. Our code will be released.
| Comments: | 13 pages, 3 figures |
| Subjects: | Machine Learning (cs.LG) |
| ACM classes: | I.2.6; I.2.7 |
| Cite as: | arXiv:2610.00367 [cs.LG] |
| (or arXiv:2610.00367v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00367 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yushuai Sun [view email]
[v1]
Wed, 30 Sep 2026 06:38:43 UTC (1,563 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org