arXiv:cs.LG· Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang·· 9 小时前AI 评分35
MaskCoFT:面向内存高效 MoE 推理的掩码协同自适应微调
MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference
AI 导读
MaskCoFT 是一种掩码协同自适应微调方法,仅用交叉熵损失同时训练路由器与专家,微调时用可学习二值掩码将每层 Top-K 路由限制到部分专家,推理时该掩码作为软先验对专家重排序,所有专家仍可被选中。
正文
Abstract:Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.34077 [cs.LG] |
| (or arXiv:2609.34077v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.34077 arXiv-issued DOI via DataCite |
Submission history
From: Junfeng Wu [view email]
[v1]
Mon, 28 Sep 2026 01:12:30 UTC (2,113 KB)
[v2]
Fri, 2 Oct 2026 14:55:49 UTC (2,113 KB)
来源:arXiv:cs.LG · arxiv.org