arXiv:cs.LG· Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi·· 3 小时前AI 评分41
自适应批处理为何有助于 LLM 预训练?无界方差视角下的解释
Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance
AI 导读
研究提出广义 BG-a 噪声模型,在方差有界(a=0)与 BG-0 噪声(a=2)之间插值,并据此推导出增长相关的 oracle 复杂度下界 Ω(ε^-(4+a)),同时通过随迭代远离初始化而增大批大小给出匹配上界。基于该理论设计的自适应批调度器在 C4 上预训练最高 1B 参数的 OLMo2 模型,在相同 token 预算下验证损失低于小批与大批训练,且迭代次数不到小批训练的 10%。
正文
Abstract:Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its success is not well understood. Analyses of stochastic optimization often assume uniformly bounded stochastic gradient variance, yet recent evidence suggests that this assumption fails in many practical nonconvex problems. The Blum--Gladyshev (BG-$0$) noise model relaxes this assumption by allowing the variance to grow quadratically with the distance from initialization, suggesting that batch size schedulers can help by controlling the variance growth during training. However, this growth can be overly conservative in practice. We empirically investigate variance growth in LLM pretraining and observe that a generalized BG model with a tunable growth exponent provides a tighter description of practical noise behavior. Motivated by this observation, we introduce the generalized BG-$a$ noise model, which interpolates between bounded variance ($a=0$) and BG-$0$ noise ($a=2$). Under $L$-smoothness, we derive an information-theoretic lower bound with growth-dependent oracle complexity $\Omega(\epsilon^{-(4+a)})$ and establish a matching upper bound in $\epsilon$-dependence by increasing the batch size as the iterates move away from initialization. Finally, we propose an adaptive batch scheduler that controls variance growth through dynamic batch size adjustments during training. In pretraining OLMo2 models of up to 1B parameters on C4, our scheduler achieves a lower validation loss than both small and large batch training under matched token budgets, while using less than 10\% of the iterations of small batch training.
| Subjects: | Machine Learning (cs.LG); Optimization and Control (math.OC) |
| Cite as: | arXiv:2610.02355 [cs.LG] |
| (or arXiv:2610.02355v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02355 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Arda Fazla [view email]
[v1]
Thu, 1 Oct 2026 18:32:34 UTC (548 KB)
来源:arXiv:cs.LG · arxiv.org