arXiv:cs.LG(机器学习,全量分类)· Hang Zou, Chao Zhang, Yuzhi Yang, Yu Tian, Samson Lasaulce, M\'erouane Debbah·· 14 小时前AI 评分37
FedFit:通过向量库参数化与量化实现 LLM 联邦微调
FedFit: Federated Fine-Tuning of LLMs via Vector-Bank Parameterization and Quantization
AI 导读
FedFit 提出一种联邦微调框架,通过不相交共享向量库参数化重建高维适配器矩阵,并用交替优化调度解决 LoRA 在联邦学习中的聚合困境。该方法结合分块量化与客户端误差反馈压缩传输向量,在 Qwen2.5 模型上困惑度与标准联邦 LoRA 方法相当,压缩比最高提升 100 倍。
正文
Abstract:Federated Learning (FL) enables privacy-preserving fine-tuning of Large Language Models (LLMs), yet the massive communication overhead remains a critical bottleneck. Furthermore, applying Low-Rank Adaptation (LoRA) in FL faces a fundamental "aggregation dilemma" between the accurate Sum-of-Products (SoP) and the communication-efficient Product-of-Sums (PoS) implementations. To tackle these challenges, we propose FedFit. First, to significantly reduce communication overhead, we introduce a disjoint shared vector-bank parameterization that reconstructs high-dimensional adapter matrices from two compact and disjoint global vector banks. Second, to address the aggregation dilemma, we devise an alternating optimization schedule. By cycling between decoupled single-bank updates (which allow for accurate aggregation) and joint updates corrected by a Residual Spectral Aggregation mechanism, we resolve the conflict between SoP and PoS. Additionally, we integrate blockwise quantization with client-side error feedback to further compress the transmitted vectors. Furthermore, we establish theoretical convergence guarantees for the proposed algorithm. Extensive experiments on Qwen2.5 models demonstrate that FedFit achieves perplexity performance comparable to standard federated LoRA methods, while providing compression ratios up to 100x higher.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.01537 [cs.LG] |
| (or arXiv:2610.01537v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01537 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hang Zou [view email]
[v1]
Thu, 1 Oct 2026 12:09:46 UTC (213 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org