跳到正文
arXiv:cs.LG· Mohammadreza Rostami, Shahriar Talebi, Solmaz S. Kia·· 5 小时前AI 评分30

ScalarFedLQR:面向线性二次调节器的标量联邦学习算法

Scalar Federated Learning for Linear Quadratic Regulator

AI 导读

ScalarFedLQR 是一种面向协作智能体 LQR 控制的通信高效联邦算法,每个智能体仅上传本地零阶梯度估计的标量投影,将单智能体上行通信从 O(d) 降至 O(1),且与策略维度无关。参与智能体越多,投影近似误差越小,可支持更大步长并实现更快的线性收敛,平均 LQR 代价线性下降。数值结果显示其性能接近全梯度联邦 LQR,通信量大幅降低,在异构但相似的智能体上同样适用。

正文

View PDF HTML (experimental)

Abstract:We propose ScalarFedLQR, a communication-efficient federated algorithm for model-free learning of a common policy in linear quadratic regulator (LQR) control of cooperative agents. The method builds on a decomposed projected gradient mechanism, in which each agent communicates only a scalar projection of a local zeroth-order gradient estimate. The server aggregates these scalar messages to reconstruct a global descent direction, reducing per-agent uplink communication from O(d) to O(1), independent of the policy dimension. Crucially, the projection-induced approximation error diminishes as the number of participating agents increases, yielding a favorable scaling law: larger fleets enable more accurate gradient recovery, admit larger stepsizes, and achieve faster linear convergence despite high dimensionality. Under standard regularity conditions for homogeneous agents, all iterates remain stabilizing and the average LQR cost decreases linearly fast. Numerical results demonstrate performance comparable to full-gradient federated LQR with substantially reduced communication, even for heterogeneous (but similar) agents.
Subjects: Systems and Control (eess.SY); Machine Learning (cs.LG)
Cite as: arXiv:2604.05088 [eess.SY]
  (or arXiv:2604.05088v2 [eess.SY] for this version)
  https://doi.org/10.48550/arXiv.2604.05088

arXiv-issued DOI via DataCite

Submission history

From: Shahriar Talebi [view email]
[v1] Mon, 6 Apr 2026 18:42:31 UTC (1,340 KB)
[v2] Thu, 1 Oct 2026 19:44:34 UTC (1,547 KB)

来源:arXiv:cs.LG · arxiv.org