arXiv:cs.LG· Zihan Fang, Qianru Wang, Haonan An, Zheng Lin, Yiqin Deng, Symeon Chatzinotas, Yuguang Fang·· 7 小时前AI 评分31
FedAlign-MoE:数据异构下移动边缘网络的联邦 MoE 对齐
Federated Mixture-of-Experts Alignment on Mobile Edge Networks under Data Heterogeneity
AI 导读
针对移动边缘设备上基于 MoE 的大语言模型联邦微调中数据异构导致的门控偏好分歧与同索引专家语义模糊问题,研究者提出联邦聚合对齐框架 FedAlign-MoE,通过一致性加权对齐路由分布、分布正则优化本地门控网络,并选择性聚合语义对齐的专家。实验显示其在 non-IID 联邦环境下收敛更快、精度更高,且计算轻量、通信高效。
正文
Abstract:The growing demand for on-device large language model (LLM) services on mobile edge devices has driven the adoption of Mixture-of-Experts (MoE) architectures, which scale model capacity with limited computation. Since fine-tuning MoE-based LLMs relies on privacy-sensitive local data, federated learning (FL) offers a natural paradigm for collaborative training without exposing raw data. However, integrating MoE-based LLM fine-tuning into FL faces two critical challenges caused by data heterogeneity across clients: (i) divergent local data distributions drive clients to develop distinct gating preferences, so direct parameter aggregation yields a one-size-fits-none global gating network; and (ii) same-indexed experts develop disparate semantic roles across devices, leading to expert semantic blurring and degraded specialization. To address these challenges, we propose FedAlign-MoE, a federated aggregation alignment framework for edge computing systems that jointly enforces routing consistency and expert semantic alignment. Specifically, FedAlign-MoE aggregates gating behaviors by aligning routing distributions through consistency weighting and optimizes local gating networks through distribution regularization, maintaining cross-client stability while preserving discriminative local gating preferences. Meanwhile, FedAlign-MoE quantifies the semantic consistency of same-indexed experts across devices and selectively aggregates semantically aligned experts, ensuring stable and specialized global experts. Extensive experiments demonstrate that FedAlign-MoE outperforms state-of-the-art benchmarks, achieving faster convergence and higher accuracy in non-IID federated environments with lightweight computation and efficient communication.
| Comments: | 15 pages, 17 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2603.21276 [cs.LG] |
| (or arXiv:2603.21276v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2603.21276 arXiv-issued DOI via DataCite |
Submission history
From: Lin Zheng [view email]
[v1]
Sun, 22 Mar 2026 15:07:39 UTC (13,187 KB)
[v2]
Tue, 6 Oct 2026 15:09:51 UTC (1,560 KB)
来源:arXiv:cs.LG · arxiv.org