跳到正文
arXiv:cs.LG· Aleksandar Armacki, Haoyuan Cai, Ali H. Sayed·· 4 小时前AI 评分31

重尾噪声下的去中心化 SGD:最优收敛率与梯度裁剪的作用

Decentralized SGD under Heavy-Tailed Noise: Optimal Convergence Rates and the Role of Gradient Clipping

AI 导读

针对重尾噪声下去中心化优化能否达到最优收敛率的问题,研究者证明裁剪去中心化 SGD(clipped DSGD)在光滑非凸目标、$p \in (1,2]$ 有界 p 阶矩噪声下,同时以高概率和期望意义达到阶最优收敛率,并实现智能体数量的线性加速。技术关键在于利用裁剪结构对共识差距的精细分析,将网络效应降为高阶项;与归一化 DSGD 可能不收敛不同,裁剪保留幅度信息,从而保证收敛性。

正文

View PDF HTML (experimental)

Abstract:Heavy-tailed noise has been widely observed in modern machine learning, motivating the use of methods like gradient clipping and normalization. While these methods are well understood in centralized settings, much less is known in decentralized ones, where applying a nonlinearity to local gradients affects both optimization and consensus. Recent works on decentralized non-convex optimization have studied both clipping and normalization under heavy-tailed noise, with clipping yielding suboptimal rates and normalization needing local momentum or mini-batches to converge. This raises the question: can a baseline decentralized method using a nonlinearity achieve optimal convergence rates under heavy-tailed noise? We answer affirmatively with clipped decentralized SGD ($\mathtt{DSGD}$). For smooth non-convex costs under bounded $p$-th moment noise, $p \in (1,2]$, we show that clipped $\mathtt{DSGD}$ achieves order-optimal rates both with high probability and in expectation. Moreover, we establish a linear speed-up in the number of agents, which, to our knowledge, has not been shown for decentralized methods with clipping. The key technical ingredient is a sharp analysis of the consensus gap that exploits the structure of clipping, relegating network effects to higher-order terms. Our results highlight an important distinction between clipping and normalization in decentralized settings: while normalized $\mathtt{DSGD}$ can fail to converge, clipping retains magnitude information, enabling $\mathtt{DSGD}$ to be convergent and order-optimal. Numerical experiments validate our theory.
Comments: 36 pages, 5 figures, 2 tables
Subjects: Optimization and Control (math.OC); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as: arXiv:2610.10527 [math.OC]
  (or arXiv:2610.10527v1 [math.OC] for this version)
  https://doi.org/10.48550/arXiv.2610.10527

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Aleksandar Armacki [view email]
[v1] Wed, 7 Oct 2026 17:57:58 UTC (1,521 KB)

来源:arXiv:cs.LG · arxiv.org