arXiv:cs.LG· Xingyu Dang, Kaiyue Wen, Sadhika Malladi·· 4 小时前AI 评分43
研究:最佳优化器取决于 batch size,Muon 缩放规则难以通用
The Best Optimizer Depends on Batch Size
AI 导读
一项研究指出,最佳优化器会随 batch size 变化,即使经过大量超参数调优,语言模型预训练的最优优化器仍会改变。研究还发现,Muon 没有一套原则性缩放规则能在不同训练设置下保持一致,这挑战了现有优化器开发与评测方式。
正文
Abstract:A plethora of new adaptive optimizers are designed to efficiently estimate and use minibatch gradient statistics to shape parameter updates, but they are typically benchmarked at a single batch size. Hyperparameter scaling rules promise to preserve performance as batch size and gradient noise change, suggesting that the best optimizer at one batch size should remain the best at another. We challenge this approach to developing and evaluating optimizers by showing: (1) no principled scaling rule for Muon works consistently across training settings, and (2) the best optimizer for language model pretraining changes with batch size even after extensive hyperparameter tuning.
| Subjects: | Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.08975 [cs.LG] |
| (or arXiv:2610.08975v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08975 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kaiyue Wen [view email]
[v1]
Tue, 6 Oct 2026 18:37:45 UTC (679 KB)
来源:arXiv:cs.LG · arxiv.org