跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Syon Mansur, Joel Zylberberg·· 14 小时前AI 评分36

增加网络宽度可让贪心逐层训练在自监督学习中媲美端到端反向传播

Increasing Width Allows Greedy Layer-wise Training to Rival End-to-End Backpropagation in Self-Supervised Learning

AI 导读

研究发现,在卷积网络的自监督学习中,网络越宽,端到端反向传播相对贪心逐层训练的优势越小;在较浅且极宽的网络中,贪心逐层训练甚至取得更高性能。对表征的分析显示,极宽贪心训练网络的表征几何结构优于端到端反向传播训练的模型,表明宽度可补偿受限的信用分配。

正文

View PDF HTML (experimental)

Abstract:End-to-end backpropagation has been the dominant mode of training in deep learning, allowing for the coordination of parameter updates across layers of a neural network. Prior studies have explored alternative -- and, in some cases, simpler -- training mechanisms, showing that they can sometimes achieve performance similar to backpropagation. However, the architectural conditions under which locally optimized networks, which avoid end-to-end backpropagation of error, can learn representations comparable to those learned through end-to-end training remain unclear. We aim to answer this question in the context of self-supervised learning, an important framework for large-scale pretraining in artificial intelligence. Here, we investigate how network width and depth affect the efficacy of greedy layer-wise and end-to-end self-supervised training in convolutional networks. We find that in wider networks, the benefits of end-to-end backpropagation over greedy layer-wise training shrink: in relatively shallow and very wide networks, we even observed higher performance in models trained with greedy layer-wise training. Subsequent analysis of the representations formed by these networks shows that very wide greedy-trained networks exhibit more favorable representational geometry than do networks trained end-to-end with backpropagation. This work shows that width can compensate for restricted credit assignment and identifies differences in representational geometry as a potential mechanism for their improved performance.
Comments: 10 pages, 5 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Neural and Evolutionary Computing (cs.NE); Neurons and Cognition (q-bio.NC)
ACM classes: I.2.6; I.5.1
Cite as: arXiv:2610.00753 [cs.LG]
  (or arXiv:2610.00753v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00753

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Syon Mansur [view email]
[v1] Wed, 30 Sep 2026 21:47:26 UTC (381 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org