arXiv:cs.LG· Yucheng Chen·· 6 小时前AI 评分48
合成基准夸大了 Forward-Forward 的可扩展性:层局部训练在真实数据上的局限
Synthetic Benchmarks Overstate Forward-Forward Scaling: Real-Data Limits of Layer-Local Training
AI 导读
研究提出 DTG-FF 方法,在九个真实数据基准上刷新 FF 系列最优成绩,包括 CIFAR-10 达 91.8%。
正文
Abstract:Forward-Forward (FF) learning [Hinton, 2022] replaces backpropagation with strictly layer-local goodness updates. Recent FF-CNN work has narrowed the gap to BP on 32x32 benchmarks, raising the question of whether layer-local training is becoming a viable alternative at realistic scale. To probe this rigorously, we develop DTG-FF -- dynamic temperature goodness, decoupled normalization, and multi-layer fusion -- as an instrument that sets FF-family state of the art across nine real-data benchmarks at the time of our submission (91.8% CIFAR-10 and an FF baseline at ImageNet-100 224x224), and use it to audit how far layer-local training actually scales.
(1) Real-data scaling. Under identical recipe and backbone, an architecture-matched BP-DeepSup baseline beats DTG-FF by 2.40/5.93 pp on CIFAR-10/CIFAR-100, and the gap widens with class count. At 224x224 the same instrument reaches only 49.4% on ImageNet-100, versus 75.6% even for a self-supervised BP-trained ResNet-50 with linear evaluation [Tian et al., 2020] -- exposing a real-data ceiling invisible at 32x32.
(2) Synthetic vs. real K-conflict. DTG-FF increasingly outperforms BP as class count K grows on synthetic teacher-student tasks, yet on real images the FF-BP gap reverses sign and widens with K. A within-dataset CIFAR-100 coarse vs. fine probe isolates label-hierarchy from image distribution: synthetic K-sweeps confound output dimensionality with fine-grained discrimination difficulty and thereby overstate FF transferability.
(3) Systems audit. FF can be implemented without storing depth-wide activations, but on commodity 8 GB hardware standard BP+gradient-accumulation reaches 4.18 GB / 157 imgs/s versus DTG-FF's 7.90 GB / 138 imgs/s, so a memory-based justification for FF at this scale is not supported under fair baselines.
| Comments: | 26 pages, 6 figures. v2: corrected bibliography (several v1 references were erroneous or nonexistent) and errors in descriptions of prior work; qualified state-of-the-art claims as of submission and added concurrent work; corrected minor factual and arithmetic errors |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE) |
| Cite as: | arXiv:2606.06539 [cs.CV] |
| (or arXiv:2606.06539v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2606.06539 arXiv-issued DOI via DataCite |
Submission history
From: Yucheng Chen [view email]
[v1]
Thu, 4 Jun 2026 04:01:01 UTC (284 KB)
[v2]
Wed, 7 Oct 2026 02:29:42 UTC (285 KB)
来源:arXiv:cs.LG · arxiv.org