跳到正文
arXiv:cs.LG· Bing Liu, Wenjie Zhou, Chengcheng Zhao, Hongtao Zhang, Boao Kong, Felix Dangel, Wu Lin·· 4 小时前AI 评分32

Bregman 散度如何塑造 Shampoo 优化器

How Bregman Divergences Shape Shampoo

AI 导读

研究提出统一 Bregman 散度框架,将 Frobenius、KL 等常见散度统一分析,揭示散度选择如何影响 Shampoo 优化器的 Kronecker 近似与预条件效果。通过梯度二阶矩的经验谱分析发现,部分散度能更好补偿有限样本下二阶矩的低估,解释了不同 Shampoo 变体的行为差异,并经 GPT-2 预训练实验验证。

正文

View PDF HTML (experimental)

Abstract:Understanding the principles behind Shampoo has recently guided the development of more effective neural network optimizers. These methods learn a preconditioner by optimizing the Frobenius or Kullback-Leibler (KL) divergence against the gradient second moment. In this work, we investigate how the choice of divergence shapes preconditioning, which remains unclear and blocks further improvements. To do so, we develop a unified Bregman divergence framework that connects all popular divergences, allowing us to study them jointly. Through empirical spectral analysis of gradient second moments, we examine how divergence choice shapes Kronecker approximation and interacts with finite-sample error in preconditioning. We find that some divergences can better compensate for finite-sample underestimation of the empirical second moment, helping explain the differing behavior of their corresponding Shampoo variants. We further validate this explanation through GPT-2 pretraining experiments. By connecting divergence choice to practical training behavior, we believe our framework provides principled guidance for understanding the foundations of, and further improving, Shampoo.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.08534 [cs.LG]
  (or arXiv:2610.08534v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.08534

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Bing Liu [view email]
[v1] Tue, 6 Oct 2026 15:25:57 UTC (9,959 KB)

来源:arXiv:cs.LG · arxiv.org