arXiv:cs.LG· Amartya Roy, Souvik Chakraborty·· 3 小时前AI 评分35
Muon 有限步 Newton–Schulz 正交化的训练损失保证
Training-Loss Guarantees for Muon with Finite-Step Newton--Schulz Orthogonalization
AI 导读
针对 Muon 收敛性分析此前只保证稳定点、且假设精确正交化的问题,研究者给出了计入动量累积与调优有限步更新的有限时间训练保证。在足够宽的两层 ReLU 网络全批量训练下,Muon 能以高概率达到任意目标经验平方损失 ε,命中时间上界为 O((1-μ)^{-1}ε^{-1/2}),所需宽度与目标精度和动量无关。数值实验中,六个宽度、五种学生初始化共 30 次运行均达到目标损失并保持核正定性。
正文
Abstract:Existing convergence analyses of Muon either assume exact orthogonalization or analyze classical Newton--Schulz polynomials, and guarantee only stationarity, so it is unresolved what Muon's five tuned Newton--Schulz steps preserve and whether that suffices to reach a prescribed neural-network training loss. We establish a finite-time training guarantee that accounts for both momentum accumulation before orthogonalization and the tuned finite-step update. For full-batch training of a sufficiently wide two-layer ReLU network with fixed random output weights and a positive-definite limiting neural tangent kernel, we prove that Muon reaches any target empirical squared loss $\varepsilon>0$ with high probability over initialization. For every momentum parameter $\mu\in[0,1)$, a target-dependent constant learning rate proportional to $(1-\mu)\sqrt{\varepsilon}$ yields a hitting-time bound of $O((1-\mu)^{-1}\varepsilon^{-1/2})$, with other problem parameters fixed. The sufficient width is independent of both target accuracy and momentum. The analysis shows that the tuned Newton--Schulz map preserves alignment with the momentum buffer while bounding the update's spectral norm. Control of gradient variation near initialization transfers this alignment to the current gradient, ensuring descent until the target is reached without requiring exact orthogonalization. Numerical experiments support these mechanisms at widths below the sufficient theoretical threshold: gradient-update alignment remains above the analytical reference, and all 30 runs across six widths and five student initializations on a fixed teacher-student dataset reach the target loss while maintaining kernel positivity.
| Comments: | 22 pages, 5 figures |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC) |
| Cite as: | arXiv:2610.03306 [cs.LG] |
| (or arXiv:2610.03306v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03306 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Amartya Roy [view email]
[v1]
Fri, 2 Oct 2026 13:45:57 UTC (314 KB)
来源:arXiv:cs.LG · arxiv.org