跳到正文
arXiv:cs.AI· Manfred M. Fischer, Joshua Pitts·· 4 小时前AI 评分32

深度 CNN 的有效深度悖论:VGG、ResNet 与 GoogLeNet 的拓扑与可训练性对比研究

The Effective Depth Paradox: Topology and Trainability in Deep CNNs

AI 导读

一项在 CIFAR-10 统一训练协议下对比 VGG、ResNet、GoogLeNet 的受控研究提出"有效深度悖论":VGG 式堆叠随有效深度增加出现精度提前饱和,而 ResNet 和 GoogLeNet 通过保持有效深度相对名义深度较低,持续从加深中获益。

正文

View PDF HTML (experimental)

Abstract:This paper presents a controlled comparative study of convolutional neural network (CNN) topology and image classification performance across the architectural families VGG, ResNet, and GoogLeNet, evaluated on CIFAR-10 under a unified training protocol. We formalize the distinction between nominal depth ($D_{\mathrm{nom}}$), the physical count of weight-bearing layers, and effective depth ($D_{\mathrm{eff}}$), an operational metric quantifying the expected length of forward information paths, extending the path-ensemble interpretation of residual networks introduced by Veit et al. (2016) into closed-form, pre-training proxies spanning sequential, residual, and multi-branch topologies. We validate this proxy against a gradient-weighted variant computed from observed backpropagation signal. Across eight representative models (VGG-11/13/16/19, ResNet-18/34/50, GoogLeNet), plain VGG-style stacks show early accuracy saturation as $D_{\mathrm{eff}}$ increases, whereas ResNet and GoogLeNet continue to benefit from added depth by keeping $D_{\mathrm{eff}}$ low relative to $D_{\mathrm{nom}}$ - a pattern we term the "Effective Depth Paradox". A pooled correlation analysis shows both $D_{\mathrm{nom}}$ and $D_{\mathrm{eff}}$ are strongly, significantly associated with accuracy (r = 0.94 and r = 0.93; both p < 0.01); given the small family-clustered sample, this alone cannot cleanly separate the two metrics, so we treat gradient-norm evidence as complementary mechanistic support rather than decisive statistical proof. We conclude that architectural topology, not layer count alone, governs trainability and scaling efficiency in deep CNNs. All claims are scoped to CIFAR-10-scale training of the three families studied; we do not claim validation at ImageNet scale or generalization to modern architectures such as EfficientNet, ConvNeXt, or Vision Transformers, which we identify as necessary future work.
Subjects: Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI)
Cite as: arXiv:2602.13298 [cs.CV]
  (or arXiv:2602.13298v4 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2602.13298

arXiv-issued DOI via DataCite

Submission history

From: Manfred M. Fischer [view email]
[v1] Mon, 9 Feb 2026 10:14:15 UTC (161 KB)
[v2] Fri, 27 Mar 2026 09:02:37 UTC (152 KB)
[v3] Fri, 8 May 2026 11:59:25 UTC (166 KB)
[v4] Fri, 2 Oct 2026 09:51:14 UTC (491 KB)

来源:arXiv:cs.AI · arxiv.org