arXiv:cs.LG(机器学习,全量分类)· Shanbin Yu, Shaoyang Guo, Haoran Zhao, Danni Yu, Ziming Liu·· 1 天前AI 评分47
NanoGPT 研究揭示「分叉」现象:数据回放下的突然过拟合
Forking: Sudden Overfitting Under Replay
AI 导读
论文研究了在 NanoGPT 自动研究中发现的泛化失败现象「分叉」(forking):数据回放时,带过度编码 n-gram 记忆分支的模型在 epoch 边界出现训练与验证损失的急剧分离。
正文
Abstract:This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily crowded tables suppress it. We also observe forking in short-budget, heavily repeated SFT and RL-like regimes. The contributions of this paper are twofold: (1) Forking reveals yet another curious phenomenon in deep learning, in addition to grokking and double descent. (2) Forking is an unexpected and unpleasant by-product of tricks proposed by autoresearch agents. While these agents produce an enormous number of results that seem useful, we should always be careful with their results.
| Comments: | 42 pages, 22 figures. Code and reproduction materials: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| MSC classes: | 68T07 |
| ACM classes: | I.2.6; I.2.7 |
| Cite as: | arXiv:2610.00394 [cs.LG] |
| (or arXiv:2610.00394v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00394 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shaoyang Guo [view email]
[v1]
Wed, 30 Sep 2026 11:40:10 UTC (6,070 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org