跳到正文
arXiv:cs.LG· Eugene Golikov, Yaroslav Gusev, Dmitry Yarotsky·· 4 小时前AI 评分33

大型线性自编码器的学习机制棱镜层级

A prism hierarchy of learning regimes in large linear autoencoders

AI 导读

该研究为大型权重绑定线性自编码器提出了一套系统性学习机制图景,证明其极端学习机制在损失展开层级上对应三棱镜的各面。研究识别出五种基本极端机制:大数据、小数据、平均场、窄隐层和自由机制,并对前四种严格推导了梯度流下的极限训练与总体动力学,证明了收敛性且与实验结果吻合。分析揭示了谱临界慢化、初始化导致偏离 PCA 最优的陷阱,以及总体损失指数收敛而训练损失仅代数松弛的机制。

正文

View PDF HTML (experimental)

Abstract:Theoretical studies of machine learning models commonly consider different limiting regimes in which the learning dynamics of gradient descent becomes theoretically tractable. It is, however, desirable to have a systematically obtained picture of qualitatively different extreme learning regimes for a particular type of models. In this paper we propose such a picture for large weight-tied linear autoencoders characterized by input and latent dimensions, initialization magnitude, and training set size. This model is nonlinear in the weights and its gradient flow does not have a general theoretical solution. We show that at the level of the formal loss-expansion hierarchy, its extreme regimes are naturally associated with faces of a triangular prism. In particular, there are five basic extreme regimes associated with the 2-faces of the prism: (1) large-data, (2) small-data, (3) mean-field, (4) narrow-latent, and (5) free. For regimes (1,2,3,4), we derive and rigorously characterize the limiting train and population dynamics under gradient flow, rigorously prove convergence to these limits, and get good agreement with experimental results. The resulting limits reveal spectral critical slowing, initialization-driven trapping away from the PCA optimum, and regimes in which population loss converges exponentially while training loss relaxes only algebraically.
Comments: 95 pages; polished and slightly extended version of a paper under review for ICLR'2027
Subjects: Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as: arXiv:2606.05335 [cs.LG]
  (or arXiv:2606.05335v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2606.05335

arXiv-issued DOI via DataCite

Submission history

From: Eugene Golikov [view email]
[v1] Wed, 3 Jun 2026 18:24:05 UTC (238 KB)
[v2] Tue, 6 Oct 2026 19:46:13 UTC (200 KB)

来源:arXiv:cs.LG · arxiv.org