arXiv:cs.LG(机器学习,全量分类)· Akshay Balsubramani·· 15 小时前AI 评分34
SGD 方法的精确信息论分析
Exact information accounting for SGD methods
AI 导读
研究者提出对随机梯度下降(SGD)及其变体的精确信息论分析,证明预条件 SGD 步是高斯贝叶斯模型的后验均值更新,其单步 regret 可分解为内在时间成本与比较器信息变化。该恒等式统一涵盖凸收敛、严格鞍点逃逸、平坦性与泛化关系、学习率调度、自适应优化器及噪声、动量、重尾、无梯度等变体。在真实网络上,该分解将经典收敛界的松弛归因于推导中被丢弃的项,并能区分达到相同训练损失的优化器。
正文
Abstract:As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its variants. We show that a preconditioned SGD step is the posterior-mean update of a Gaussian Bayes model, and that its one-step regret splits into an intrinsic-time cost and a change in comparator information. The split extends to an identity for the objective itself. Convex convergence, strict-saddle-point escape, the link between flatness and generalization, the standard learning-rate schedules, adaptive optimizers, and the noisy, momentum, heavy-tailed, and gradient-free variants of SGD each correspond to a term or a special case of this identity. We measure its terms on synthetic and real training runs. On real networks it attributes the slack of classical convergence bounds to the terms their derivations drop and separates optimizers that reach the same training loss. That separation follows the number and consistency of their steps. Its relation to which of them generalizes better differs between networks. For gradient-free SGD the identity determines how a curvature preconditioner should enter the update. The sharpness-based generalization certificate it yields, with a data-independent isotropic prior, is vacuous at network scale unless the curvature spectrum is nearly flat across all parameters.
| Subjects: | Machine Learning (cs.LG); Information Theory (cs.IT); Optimization and Control (math.OC); Machine Learning (stat.ML) |
| MSC classes: | 68T05, 90C15, 94A17 |
| Cite as: | arXiv:2610.00446 [cs.LG] |
| (or arXiv:2610.00446v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00446 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Akshay Balsubramani [view email]
[v1]
Wed, 30 Sep 2026 17:53:04 UTC (330 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org