跳到正文
arXiv:cs.LG· Kwanhee Lee, Namhoon Lee, Dan Alistarh·· 3 小时前AI 评分33

面向万亿级 MoE 的硬件原生联合稀疏量化框架

Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts

AI 导读

研究者提出一套软硬件协同设计框架,将 MoE 专家权重压缩为硬件原生低精度稀疏表示并在 SpTC 上加速执行。在 30B 至 1T 参数的 MoE 模型上,该方法将联合稀疏量化精度提升最多 4.35 个百分点,同时保留原模型 96.09% 的性能。

正文

View PDF HTML (experimental)

Abstract:Mixture-of-Experts (MoE) architectures allow frontier language models to scale to trillions of parameters, but their deployment is constrained by massive memory footprints and memory-bandwidth limitations. Although modern accelerators provide Sparse Tensor Cores (SpTCs) that reduce weight storage and increase throughput through low-precision semi-structured sparsity, exploiting them for MoEs remains challenging because of substantial model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware-software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes discrete semi-structured support selection through continuous reparameterization, enabling differentiable joint optimization with quantized weights under a router-weighted reconstruction objective and scalable expert-parallel compression. Systemically, we develop a custom grouped sparse GEMM kernel tailored to low-precision sparse MoE inference on SpTCs. Across MoE models ranging from 30 billion to one trillion parameters, our framework improves state-of-the-art joint sparse-quantization accuracy by up to 4.35 percentage points while preserving 96.09% of the original model's performance. On NVIDIA B200 GPUs, our kernel outperforms the vendor baseline by up to $1.65\times$, increasing serving throughput by $1.18\times$ and reducing end-to-end latency by up to $4.03\times$. These results establish hardware-software co-design as a practical path toward scalable and efficient MoE deployment.
Subjects: Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.02241 [cs.AR]
  (or arXiv:2610.02241v1 [cs.AR] for this version)
  https://doi.org/10.48550/arXiv.2610.02241

arXiv-issued DOI via DataCite

Submission history

From: Kwanhee Lee [view email]
[v1] Tue, 29 Sep 2026 07:58:06 UTC (950 KB)

来源:arXiv:cs.LG · arxiv.org