arXiv:cs.LG· Kota Maejima, Takayuki Nishio, Asato Yamazaki, Yuko Hara-Azumi·· 4 小时前AI 评分32
Tram-FL:通过模型顺序循环降低去中心化联邦学习的通信与计算成本
Tram-FL: Reducing Communication and Computation Costs through Sequential Model Circulation in Decentralized Federated Learning
AI 导读
研究者提出 Tram-FL(Traveling Model Training Mechanism),通过让单个模型在节点间顺序循环训练来实现去中心化联邦学习,而非让每个客户端各自维护模型副本。该方法针对模型循环中的训练调度问题,考虑循环路径与更新迭代分配,并结合量化动量在减少循环次数的同时控制每次传输的通信负载。实验显示该算法在 non-IID 数据下仍能以更低的通信和计算成本收敛到全局模型。
正文
Abstract:Conventional decentralized federated learning (DFL) often focuses on clients, with each client maintaining a model copy, performing updates individually, and undertaking model exchange and integration. While fully leveraging computational resources can shorten training times, it can also lead to significant computational and communication waste. This is especially pronounced with non-independent and identically distributed (non-IID) data, where achieving high model accuracy demands extra resources. This research shifts focus to the model itself, aiming to realize DFL with minimal computation and communication costs. To this end, we propose Tram-FL (Traveling Model Training Mechanism for Decentralized Federated Learning), a mechanism designed to efficiently address these challenges. It sequentially trains a single model by circulating it among nodes. We address the training scheduling problem in model circulation-based training, specifically determining which nodes should update the model and the number of updates to perform. This is approached by considering the model's circulation route and update iteration allocation, for which we propose simple yet effective methods. Additionally, with quantized momentum, Tram-FL achieves high accuracy with fewer model circulations while controlling communication load per transmission. Experimental results show that the proposed algorithm, even with non-IID data, converges to a global model with reduced communication and computation.
| Comments: | 12 pages, 7 figures, 5 tables. This work has been submitted to the IEEE for possible publication |
| Subjects: | Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Networking and Internet Architecture (cs.NI) |
| Cite as: | arXiv:2610.07859 [cs.LG] |
| (or arXiv:2610.07859v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07859 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Takayuki Nishio [view email]
[v1]
Tue, 6 Oct 2026 07:01:01 UTC (523 KB)
来源:arXiv:cs.LG · arxiv.org