跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Junghoon Seo·· 15 小时前AI 评分29

常步长 SGD 不变分布的精确高斯近似

Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD

AI 导读

研究证明,在由外生一致遍历 Markov 链生成有界加性噪声时,常步长 SGD 的不变分布可用高斯分布精确近似:对光滑强凸且 Hessian 满足 Lipschitz 的目标函数,中心化迭代除以步长平方根后,与极限高斯分布的 1-Wasserstein 距离为 O(√α)。

正文

View PDF HTML (experimental)

Abstract:We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is $O(\sqrt{\alpha})$-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.
Comments: To be presented at 2026 NeurIPS workshop on "Optimization for Machine Learning" (OPT2026)
Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC)
Cite as: arXiv:2609.39144 [cs.LG]
  (or arXiv:2609.39144v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.39144

arXiv-issued DOI via DataCite

Submission history

From: Junghoon Seo [view email]
[v1] Wed, 30 Sep 2026 07:11:02 UTC (29 KB)
[v2] Thu, 1 Oct 2026 07:48:40 UTC (29 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org