跳到正文
arXiv:cs.LG· Yossi Arjevani·· 4 小时前AI 评分38

深度学习 Hessian 特征值:对称性的起源与破缺

Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking

AI 导读

该论文提出,深度学习训练后模型 Hessian 谱中特征值聚簇、近零大块与少量孤立离群点的现象,源于原配置对某个隐藏的高度对称参考配置的偏离。修改架构、数据分布或参数度量可暴露该参考配置,其 Hessian 具有权重对称性无法解释的丰富不变性,从而强制产生高维核与大重数特征值。

正文

View PDF HTML (experimental)

Abstract:Hessian spectra at trained models in deep learning exhibit a persistent pattern: eigenvalues organize into distinct clusters, including a large bulk near zero and a few isolated outliers. This paper shows that a natural account of these spectral phenomena emerges when the original setting is understood as a departure from a nearby, otherwise hidden, highly symmetric reference.
Modifications, including changes to the architecture, data distribution, or parameter metric, expose a nearby reference configuration whose Hessian exhibits rich invariances-ones not accounted for by weight symmetries. There, symmetry enables a precise description of the spectra, forcing high-dimensional kernels and eigenvalues of large multiplicity. Returning to the original configuration breaks the Hessian symmetry and thereby produces the observed hierarchy of clusters and outliers.
The framework is developed in some generality, with a detailed analysis of three-layer ReLU networks and applications to convolutional, graph, and transformer models, as well as to the NTK. The same mechanism is further shown to yield analogous spectral structures in layerwise Hessians and the Gauss-Newton matrix.
Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as: arXiv:2610.09919 [cs.LG]
  (or arXiv:2610.09919v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09919

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yossi Arjevani [view email]
[v1] Wed, 7 Oct 2026 12:08:48 UTC (1,857 KB)

来源:arXiv:cs.LG · arxiv.org