跳到正文
arXiv:cs.LG· Deborah Oliveira, Elliot Paquette·· 4 小时前AI 评分41

随机特征网络中的二次弱到强泛化:基于随机矩阵理论的分析

Quadratic Weak-to-Strong Generalization in Random Feature Networks via Random Matrix Theory

AI 导读

研究在两层随机特征网络中发现,当模型强度由宽度决定时,用弱教师模型标签训练的强学生模型可实现二次弱到强泛化:学生误差为教师误差的平方。在 ReLU 激活和纯球谐目标下,基于高斯普适性假设得到该渐进结果,达到 Medvedev 等人(2025)的通用下界。研究还刻画了更一般停止时间和多谐波阶目标下二次、非二次与无改进之间的转变。

正文

View PDF HTML (experimental)

Abstract:Weak-to-strong generalization is the phenomenon where a strong student model trained with labels produced by a weak teacher model is able to generalize better than the teacher. In this paper, we study this phenomenon in two-layer random feature networks where the model strength is determined by its width. Using tools from random matrix theory, we derive deterministic equivalents for the population errors of an optimally trained teacher and a student trained with gradient flow. For ReLU activation and a pure spherical harmonic target, we obtain sharp asymptotics under a Gaussian universality assumption, showing a quadratic improvement: the student error scales as the square of the teacher error. These results attain the general lower bound of Medvedev at al (2025). We also analyze how the student behaves under more general stopping times and targets supported on multiple harmonic degrees, characterizing the regimes in which weak-to-strong generalization occurs and identifying the transition between quadratic, non-quadratic, and no improvement.
Subjects: Machine Learning (stat.ML); Machine Learning (cs.LG); Probability (math.PR)
Cite as: arXiv:2610.09044 [stat.ML]
  (or arXiv:2610.09044v1 [stat.ML] for this version)
  https://doi.org/10.48550/arXiv.2610.09044

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Elliot Paquette [view email]
[v1] Tue, 6 Oct 2026 19:46:36 UTC (251 KB)

来源:arXiv:cs.LG · arxiv.org