arXiv:cs.LG(机器学习,全量分类)· Junkang Liu·· 14 小时前AI 评分31
FedMIX-P:混合局部与全局预条件器的联邦视觉与语言模型训练
FedMIX-P: Mixing Local and Global Preconditioners for Federated Vision and Language Model Training
AI 导读
FedMIX-P 在联邦学习的每个本地步骤中混合共享与局部预条件器,将均方算子失配降低 λ² 倍。该方法在随机梯度和部分参与下取得 O(R^{-1/2}) 的稳定性界,且不要求局部预条件器相互收敛。在 SOAP、Sophia 和 Muon 变体上的视觉与语言任务实验中,精度最高提升 19.47 个百分点,60M–350M 语言模型验证损失更低。
正文
Abstract:Adaptive preconditioners accelerate model training, but heterogeneous client geometries can bias federated updates even when gradients are evaluated at the same model. Round-start synchronization alone cannot prevent this mismatch from reappearing during local training. We propose \texttt{FedMIX-P}, which mixes shared and local preconditioners at every local step, retaining local adaptation while reducing mean-squared operator mismatch by a factor of $\lambda^2$. For smooth nonconvex objectives with stochastic gradients and partial participation, we establish an $O(R^{-1/2})$ stationarity bound using suitable stepsizes and a horizon-dependent mixing weight, without requiring local preconditioners to converge to one another. A two-client counterexample shows that fixed positive mixing can preserve a nonstationary fixed point. The theory covers bounded linear symmetric positive-definite preconditioners. Experiments with SOAP, Sophia, and Muon variants across vision and language tasks show improvements over corresponding local optimizers, including accuracy gains of up to $19.47$ percentage points and lower validation loss for 60M--350M language models. Full nonlinear and momentum-based updates require separate analysis.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.01515 [cs.LG] |
| (or arXiv:2610.01515v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01515 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Junkang Liu [view email]
[v1]
Thu, 1 Oct 2026 11:50:34 UTC (71 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org