arXiv:cs.LG· Vishrut Goyal, Rohan Ramkumar·· 4 小时前
Spectrally Targeted Muon:改写 Muon 优化器的正交化策略
Spectrally Targeted Muon
AI 导读
研究者提出 Spectrally Targeted Muon,只对阈值 τ 以上或以下的奇异值做正交化,使 τ 变化可在归一化 SGD 与 Muon 之间插值,并用 Newton-Schulz 迭代在移位 Gram 矩阵上计算投影、无需 SVD。
正文
Abstract:The Muon optimizer orthogonalizes each update matrix, setting all of its singular values to one, and has proven highly effective for training large language models. It remains unclear, however, whether this success comes from amplifying small singular directions that gradient descent neglects or from suppressing large, degenerate directions that disrupt training. We introduce Spectrally Targeted Muon, which orthogonalizes only the singular values above or below a threshold $\tau$, so that varying $\tau$ interpolates between normalized SGD and Muon. It isolates the relevant singular subspaces with projections computed by Newton-Schulz iteration on a shifted Gram matrix, so no SVD is needed. We evaluate these variants on the CIFAR-10 and NanoGPT speedruns, tracking the effective rank of gradient, update, and weight matrices and a new metric, the alignment of updates with the tangent space of the weight matrix's isospectral manifold. We find three things. First, the small singular values of the momentum are not noise. Orthogonalizing everything except the few largest singular values of each matrix nearly matches Muon while touching only a small fraction of the momentum, whereas orthogonalizing only the top falls well short even though it holds almost all of it. On language models every momentum singular value is far below one, so targeted orthogonalization can only amplify, and Muon wins by making directions that are too small to train on at their raw scale trainable. Second, shrinking the largest singular values is what keeps the parameter spectrum flat; this is cheap, and it is not what drives the loss. Third, AdamW differs from Muon mainly in how slowly it builds structure, which explains its slower start and why warmup helps AdamW but only hurts Muon.
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10965 [cs.LG] |
| (or arXiv:2610.10965v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10965 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vishrut Goyal [view email]
[v1]
Wed, 7 Oct 2026 22:35:28 UTC (328 KB)
来源:arXiv:cs.LG · arxiv.org