跳到正文
arXiv:cs.AI· Pengxin Guo, Shuang Zeng, Zonggen Li, Weiying Zheng, Mengting Liu, Liangqiong Qu·· 3 小时前

Fed-GRPO:奖励信号驱动的联邦群体相对策略优化

Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization

AI 导读

Fed-GRPO 是一个联邦 GRPO 训练框架,无需共享原始数据即可协作训练 LLM 推理能力,并利用 GRPO 训练中天然产生的奖励统计作为零成本信号。它包含信号加权聚合、全局奖励校准和自适应稀疏通信三种机制,在数学推理任务上性能优于所有联邦方法、明显超过 FedAvg 并接近集中式训练水平,同时无损降低通信 32 倍,在紧张带宽预算下支持最高 621 倍压缩且仅有轻微精度下降。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by $32\times$ and supporting up to $621\times$ compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at this https URL.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11502 [cs.AI]
  (or arXiv:2610.11502v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11502

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Pengxin Guo [view email]
[v1] Thu, 8 Oct 2026 08:40:41 UTC (364 KB)

来源:arXiv:cs.AI · arxiv.org