跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Sahel Torkamani, Henry Gouk, Rik Sarkar·· 14 小时前AI 评分28

神经网络泛化研究:用 magnitude potential 度量表征与 grokking

Generalization in Neural Networks Through the Lens of Magnitude Potential

AI 导读

研究者提出 magnitude potential——基于度量 magnitude 理论、衡量任意点被给定集合表征程度的量,并将其用于分析神经网络泛化。在 logit 层计算某类与全数据的 magnitude potential 比值,该比值与 Feldman 记忆分数相关,聚合后可检测决策边界结构变化,并为模算术中的 grokking 提供几何指标。即使神经坍缩被显式抑制,该比值仍具信息量。

正文

View PDF HTML (experimental)

Abstract:Explaining generalization and training dynamics in neural networks remains a challenge, and various approaches have been developed to study different aspects of these phenomena. In this paper, we introduce the idea of {\em magnitude potential} -- a quantity based on the theory of metric magnitude -- that reflects how well an arbitrary point is represented by a given set. We find that this basic quantity can be applied to examine various features in neural generalization. The ratio between the magnitude potential with respect to a class and with respect to the entire data, computed at the logit layer, is informative of the representation of the point. In experiments, these ratios for individual training points are found to be correlated with the Feldman memorization scores. Magnitude potential ratios aggregated across points detect structural changes in the decision boundaries and provide a geometric indicator of grokking in modular arithmetic. Although the magnitude potential ratio and neural collapse are both closely associated with intra-class and inter-class geometric structure, the magnitude potential ratio remains informative even when neural collapse is explicitly suppressed.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01633 [cs.LG]
  (or arXiv:2610.01633v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01633

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sahel Torkamani [view email]
[v1] Thu, 1 Oct 2026 13:03:16 UTC (5,606 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org