arXiv:cs.LG(机器学习,全量分类)· Gil Kur, Ileana Rugina, Cl\'ementine Carla Juliette Domin\'e, Marco Mondelli·· 5 小时前AI 评分36
从最小范数插值视角理解 Grokking:稀疏正则化如何决定延迟泛化
Grokking through the Lens of Minimum-Norm Interpolation
AI 导读
一项统计理论从最小范数插值视角刻画了 Grokking 中正则化几何与信号稀疏性如何决定泛化。研究在强过参数化无噪声问题上证明了零一泛化定律,并构造出一族凸范数,其插值器在训练误差保持为 0 的同时从全零预测器的平凡风险过渡到精确恢复。实验在对角线性网络和模运算 Transformer 上验证了理论预测,并揭示最小范数插值存在统计不稳定性:正则化强度的微小扰动即可导致泛化性能剧变。
正文
Abstract:Grokking shows that fitting the training data and learning the underlying signal can occur at very different stages. However, existing theories offer limited quantitative insight into how this delayed generalization depends on inductive bias and signal structure. Our work addresses the gap by developing a statistical theory that characterizes how regularization geometry and signal sparsity govern generalization near interpolation. In particular, we focus on the prototypical setting of high-dimensional regression and identify regimes in which sparsity-promoting regularization makes exact interpolation much more accurate than approximate fitting. In strongly overparameterized noiseless problems, we prove a zero--one generalization law and construct a family of convex norms whose interpolators transition from the trivial risk of the all-zero predictor to exact recovery, while keeping the training error equal to $0$. Furthermore, when feature dimension and sample size are proportional, we provide a precise characterization of training and generalization errors along $\ell_r$-regularization paths. This in turn allows us to quantify the generalization gain that remains near interpolation: we show that this gain increases as the norm becomes more sparsity-promoting and as the target becomes sparser, with a sharp drop in generalization reached for noiseless data and $\ell_1$ regularization. Experiments on diagonal linear networks and transformers trained on modular arithmetic demonstrate the generality of our theoretical predictions. Finally, beyond grokking, our work reveals a statistical instability in minimum-norm interpolation: small perturbations in the regularization strength can lead to drastically different generalization, while preserving small training error.
| Comments: | Submitted |
| Subjects: | Machine Learning (cs.LG); Statistics Theory (math.ST); Machine Learning (stat.ML) |
| Cite as: | arXiv:2609.38453 [cs.LG] |
| (or arXiv:2609.38453v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38453 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Gil Kur [view email]
[v1]
Tue, 29 Sep 2026 19:44:32 UTC (374 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org