跳到正文
arXiv:cs.AI· Gabriel U. Talasso, Meghdad Kurmanji, Allan M. de Souza, Nicholas D. Lane, Leandro A. Villas·· 3 小时前

FedRouter:面向任务的语言模型个性化联邦微调

Task-Centric Personalized Federated Fine-Tuning of Language Models

AI 导读

FedRouter 是一种基于聚类的个性化联邦学习方法,为每个任务而非每个客户端构建专用模型,通过局部与全局两种聚类机制将 adapter 与任务关联,并用评估路由器把测试样本路由到最佳 adapter。在多任务数据集上,该方法在任务干扰场景下相对提升最高 6.1%,在泛化评估中相对提升最高 136%。

正文

View PDF HTML (experimental)

Abstract:Federated Learning (FL) has emerged as a promising technique for training language models on distributed and private datasets of diverse tasks. However, aggregating models trained on heterogeneous tasks often degrades the overall performance of individual clients. To address this issue, Personalized FL (pFL) aims to create models tailored for each client's data distribution. Although these approaches improve local performance, they usually lack robustness in two aspects: (i) generalization: when clients must make predictions on unseen tasks, or face changes in their data distributions, and (ii) intra-client tasks interference: when a single client's data contains multiple distributions that may interfere with each other during local training. To tackle these two challenges, we propose FedRouter, a clustering-based pFL that builds specialized models for each task rather than for each client. FedRouter uses adapters to personalize models by employing two clustering mechanisms to associate adapters with specific tasks. A local clustering that associate adapters with task data samples and a global one that associates similar adapters from different clients to construct task-centric personalized models. Additionally, we propose an evaluation router mechanism that routes test samples to the best adapter based on the created clusters. Experiments comparing our method with existing approaches across a multitask dataset, FedRouter demonstrate strong resilience in these challenging scenarios performing up to 6.1% relatively better under tasks interference and up to 136% relative improvement under generalization evaluation.
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as: arXiv:2604.00050 [cs.LG]
  (or arXiv:2604.00050v3 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2604.00050

arXiv-issued DOI via DataCite

Submission history

From: Gabriel Ukstin Talasso [view email]
[v1] Mon, 30 Mar 2026 20:01:53 UTC (724 KB)
[v2] Sun, 5 Apr 2026 19:47:01 UTC (724 KB)
[v3] Thu, 8 Oct 2026 15:44:51 UTC (724 KB)

来源:arXiv:cs.AI · arxiv.org