arXiv:cs.LG· Meher Chaitanya, Tianyi Zhou, Aristides Gionis·· 4 小时前AI 评分40
Muon 是否需要细粒度谱塑形?两波段重加权即可媲美 Freon
Does Muon Need Fine-Grained Spectral Shaping?
AI 导读
研究提出 BulkBoost 两波段谱重加权框架,用 split-minibatch 梯度差校准 Marchenko–Pastur 噪声边界,将奇异模分为边界以下的 bulk 与以上的 spike 并共享同一增益重新加权。
正文
Abstract:Muon combines current and past gradients into matrix momentum. For $M=U\Sigma V^\top$, the idealized polar update $Q=UV^\top$ gives every singular direction the same weight. We refer to this as the flat profile. Several recent optimizers replace this flat profile with fine-grained spectral maps that give each direction its own gain. We ask how much of this spectral detail a Muon update needs. Our spectral diagnostics show that approximately $94$--$97\%$ of measured singular modes lie below an estimated noise edge, yet collectively align positively with a reference gradient.
We introduce BulkBoost, a two-band spectral reweighting framework with fixed-rank and noise-calibrated variants. The latter uses split-minibatch gradient differences to calibrate a Marchenko--Pastur reference edge for Muon's Nesterov input, separating the bulk below the edge from the spikes above it. Both variants increase the bulk's relative weight through one shared gain while preserving the Frobenius norm of each matrix's unreweighted direction. For a fixed partition, our theory gives the first-order condition under which moving weight toward the bulk lowers the loss. It also quantifies the fraction of the maximal first-order improvement rate, over all per-mode reallocations, that two bands can capture. Across 30 continued-pretraining settings spanning Pythia-14M to 410M and six corpora, two-band reweighting is competitive with the fine-grained power-law profile of Freon and outperforms Spectra. Measured against Muon's flat profile, Freon reduces final loss by $0.022\%$ of the pre-adaptation loss on average, whereas the two-band variants achieve reductions of $0.073$--$0.147\%$. These observations suggest that useful departures from the flat profile are surprisingly low-dimensional: a single bulk-to-spike gain captures at least as much benefit as the fine-grained spectral profiles.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Machine Learning (stat.ML) |
| Cite as: | arXiv:2610.07497 [cs.AI] |
| (or arXiv:2610.07497v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07497 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Meher Chaitanya Pindiprolu [view email]
[v1]
Mon, 5 Oct 2026 23:02:41 UTC (1,152 KB)
来源:arXiv:cs.LG · arxiv.org