arXiv:cs.LG(机器学习,全量分类)· Junghoon Seo·· 15 小时前AI 评分29
常步长 SGD 不变分布的精确高斯近似
Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD
AI 导读
研究证明,在由外生一致遍历 Markov 链生成有界加性噪声时,常步长 SGD 的不变分布可用高斯分布精确近似:对光滑强凸且 Hessian 满足 Lipschitz 的目标函数,中心化迭代除以步长平方根后,与极限高斯分布的 1-Wasserstein 距离为 O(√α)。
正文
Abstract:We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is $O(\sqrt{\alpha})$-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.
| Comments: | To be presented at 2026 NeurIPS workshop on "Optimization for Machine Learning" (OPT2026) |
| Subjects: | Machine Learning (cs.LG); Optimization and Control (math.OC) |
| Cite as: | arXiv:2609.39144 [cs.LG] |
| (or arXiv:2609.39144v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.39144 arXiv-issued DOI via DataCite |
Submission history
From: Junghoon Seo [view email]
[v1]
Wed, 30 Sep 2026 07:11:02 UTC (29 KB)
[v2]
Thu, 1 Oct 2026 07:48:40 UTC (29 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org