arXiv:cs.AI· Sungyoon Kim, Kaan Ozkara, Youngsuk Park·· 3 小时前
用链式 LMO 优化大语言模型:TensorChain 提出并在 Qwen3 预训练上超越 Muon
Optimizing Large Language Models with Chained LMOs
AI 导读
研究提出 chained LMOs 框架,将 Muon 等由多重矩阵归一化组合而成的优化器统一为线性最小化 oracle 的组合形式。基于该框架提出的优化器 TensorChain 跨层堆叠兼容权重矩阵并对 3d 张量沿各轴归一化,在 Qwen3 0.6B 和 1.7B 预训练中平均 token 效率优于所有链式基线,在匹配验证损失下比 Muon 平均节省 9.6% token。
正文
Abstract:Muon has motivated a growing family of optimizers that compose multiple matrix normalizations, but these methods remain fragmented and lack a unified perspective. We introduce chained linear minimization oracles (chained LMOs), which cast these methods as compositions of LMOs. Despite their empirical success, many chains fall outside the standard LMO framework and can diverge on smooth convex objectives. To explain why composition can nevertheless help, we turn to linear associative memory and show that chaining can improve over Muon under anisotropic embeddings. Empirically, we propose TensorChain, a novel optimizer within the framework that stacks compatible weight matrices across different layers and normalizes the 3d tensor across its axes. In Qwen3 0.6B and 1.7B pretraining, TensorChain outperforms all chained baselines in average token efficiency, with average token savings of 9.6% over Muon at matched validation loss.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10975 [cs.LG] |
| (or arXiv:2610.10975v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10975 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sungyoon Kim [view email]
[v1]
Wed, 7 Oct 2026 22:56:55 UTC (3,151 KB)
来源:arXiv:cs.AI · arxiv.org