跳到正文
arXiv:cs.LG· Ruoyu Zhao, Mingxuan Zhang, Jianbo Dai, Jiaqi Wu, Chenyu Zhu, Tong Che·· 3 小时前

稀有门分歧会限制可塑性:梯度流何时会误判有限批 SGD

Rare Gate Disagreements Can Limit Plasticity: When Gradient Flow Mispredicts Finite-Batch SGD

AI 导读

研究者发现,群体梯度流会定性上误判有限批 SGD 的行为,机制源于稀有 ReLU 门分歧。在两单元 ReLU 回归中,源任务训练时间 T 后,大批量在线 SGD 在 T≳log(b/η) 时以高概率长期无法恢复目标任务,而小步长 SGD 仍能恢复,说明失败需小步长与长预训练的共同极限。恢复所需目标样本数满足 Nb≳e^{λT}、批量 b≳η e^{λT},模拟中恢复近似取决于分歧预算 bδ/η。

正文

View PDF HTML (experimental)

Abstract:Population gradient flow is a common tool for reasoning about how neural networks adapt, including after pretraining. We show that it can mispredict finite-batch stochastic gradient descent (SGD) qualitatively, and we trace the discrepancy to a specific mechanism. In a two-unit ReLU regression, a source task drives the two neurons toward positive proportionality and a target task rewards separating them. After source training for time $T$, gradient flow recovers on the target in time linear in $T$. Online SGD with batch size $b$ and step size $\eta$ in both phases instead fails with high probability throughout a horizon of order $e^{c/\eta}$ once $T \gtrsim \log(b/\eta)$, uniformly on an explicit set of initializations with Gaussian probability above one percent. For each fixed $T$, small-step SGD still recovers, so the failure requires the joint limit of small steps and long pretraining. At the target clone, the population instability is carried entirely by inputs on which the two ReLU gates disagree. For units at angle $\delta$ these inputs form a wedge of probability $\delta/\pi$, and weight decay shrinks the angle exponentially during pretraining. On every other input both units receive the same random linear update, which contracts their separation in conditional expectation. Bounding the cumulative probability of sampling the wedge along the exact online recursion, without a diffusion approximation, shows that recovery with fixed probability from an identical source-gradient-flow checkpoint, within $e^{c/\eta}$ updates, requires $Nb \gtrsim e^{\lambda T}$ target samples and batch size $b \gtrsim \eta e^{\lambda T}$, where $N$ counts updates and $\lambda$ is the weight decay. In simulations, recovery is approximately a function of the disagreement budget $b\delta/\eta$ and saturates in the horizon.
Comments: 25 pages, 5 figures. Tong Che leads the project
Subjects: Machine Learning (cs.LG); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as: arXiv:2610.11475 [cs.LG]
  (or arXiv:2610.11475v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.11475

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ruoyu Zhao [view email]
[v1] Thu, 8 Oct 2026 08:22:35 UTC (452 KB)

来源:arXiv:cs.LG · arxiv.org