arXiv:cs.LG(机器学习,全量分类)· S. Aaron McClendon, Jorge Gallego-Feliciano, Antonios Saravanos·· 14 小时前AI 评分43
Transformer 大规模激活的生命周期:随机诞生、权重衰减驱动的增长与竞争性整合
The Life Cycle of a Massive Activation: Stochastic Birth, Weight-Decay-Driven Growth, and Competitive Consolidation
AI 导读
研究追踪了 Transformer 中大规模激活(massive activation)的训练轨迹:承载 attention sink 的通道随随机种子变化,但在每次训练早期即稳定,随后周围通道侵蚀、sink 集中于少数冗余载体。
正文
Abstract:Massive activations, residual-stream coordinates with magnitudes far larger than typical activations, are associated with attention sinks in transformers, but how their scale is regulated during training remains incompletely understood. Combining training-trajectory analyses and controlled interventions, we trace their emergence, growth, and consolidation. Sink-carrying channels vary across random seeds but stabilize early within each run. Over longer training, surrounding channels erode and the sink concentrates onto a few redundant carriers. Across ablations, gradient attenuation follows the sink token's collective root-mean-square magnitude rather than any single channel, making collective scale central to understanding their effects. Our central result is that weight decay causally controls the turnover of global activation scale. In controlled continuations, removing decay near the peak allows this scale to keep rising, whereas retaining it produces decline even at constant learning rate. We develop a balance model for the rise and peak of massive-activation magnitude, in which AdamW-preconditioned growth opposes weight decay. Sweeping the decay coefficient $\lambda$ shifts peak timing approximately log-linearly and yields peak magnitudes scaling approximately as $\lambda^{-1/2}$, consistent with this balance. Optimizer measurements further show that preconditioning sustains the large-channel cohort against decay even when raw maintaining forces are too small to do so. Together, these findings connect the observed life cycle to scale-regulating training dynamics and establish weight decay as a training-time lever on activation magnitude.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00423 [cs.LG] |
| (or arXiv:2610.00423v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00423 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Steven McClendon [view email]
[v1]
Wed, 30 Sep 2026 15:28:07 UTC (16,997 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org