arXiv:cs.LG· Thibaut Boissin (IRIT), Thomas Massena (IRIT, DTIPG - SNCF, UT3), Mathieu Serrurier (IRIT), Franck Mamalet·· 4 小时前AI 评分39
Muon 对卷积在理论上不成立,但在实验中有效
Muon Is Theoretically Wrong For Convolutions, But Empirically Effective
AI 导读
研究指出 Muon 优化器将四维卷积核 reshape 成矩阵的做法破坏了其理论基础,为此提出在卷积算子几何中直接建模的 Convolutional Newton-Schulz(Conv-NS),该方法在保持核支撑的同时逼近极因子。
正文
Abstract:Muon, an optimizer known for its efficiency, has a clear interpretation for matrix-valued updates, but convolutional kernels are stored as four-dimensional tensors. Standard implementations reshape these tensors into matrices, a shortcut which breaks the theoretical understanding behind Muon. To investigate this, we formalize the corresponding optimization objective directly in convolutional operator geometry and introduce Convolutional Newton-Schulz (Conv-NS), which approximates the polar factor in this geometry while preserving kernel support. When applied in fast training experiments, Conv-NS and reshape-based Muon are both computationally efficient and achieve comparable accuracy on CIFAR-10 and ImageNet classification tasks. However, as one could expect a theoretically aligned Conv-NS to outperform reshape-based Muon, we investigate this mismatch between practice and theoretical understanding, with the hypothesis that exact convolutional orthogonalization may overconstrain updates. These findings highlight Muon's strong practical performance while opening directions for its further development on convolutions. Our code is publicly available at \href{this https URL}{github conv-muon}.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07103 [cs.LG] |
| (or arXiv:2610.07103v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07103 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Franck MAMALET [view email] [via CCSD proxy]
[v1]
Mon, 5 Oct 2026 14:57:01 UTC (877 KB)
来源:arXiv:cs.LG · arxiv.org