arXiv:cs.LG(机器学习,全量分类)· Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu·· 19 小时前AI 评分43
LOOM:可扩展的循环 MoE 大模型训练方案,稳定支持 9-12 次循环
Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts
AI 导读
研究者提出 LOOM,通过缩放残差更新、每轮重注入输入嵌入和逐循环路由器,解决循环 MoE 的深度诅咒与专家选择坍缩问题,将稳定循环扩展至 9-12 次。
正文
Abstract:Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that the gains from additional iterations diminish quickly and can even turn into degradation, so the extra FLOPs spent on looping yield little substantial improvement. Consequently, prior work typically settles on two loops. We identify two main obstacles to scaling looped MoE. First, looping inherits and amplifies the curse of depth: hidden-state variance grows with each iteration as residual updates accumulate, which destabilizes deep recurrence and causes representations to drift. Second, looped MoE suffers from expert selection collapse: routers repeatedly select the same experts across loops, so extra iterations add computation without adding computational diversity. Guided by this diagnosis, we propose LOOM, built on a single principle: each loop should contribute new computation while keeping the recurrent state stable. LOOM stabilizes recurrence by scaling residual updates to bound variance growth and re-injecting the input embedding at every loop, and diversifies it through per-loop routers that engage different experts and a Looping Residual that carries earlier outputs forward. Experiments across 100M-1.7B models show stable scaling to 9-12 loops. Under near-iso-FLOP, the 700M model performs best at 5 loops, reducing perplexity from 18.36 to 16.54 and improving average zero-shot accuracy from 38.84% to 39.53% over the non-looped baseline. Without FLOP matching, the 1.7B model trained on 60B tokens peaks at 9 loops, reducing perplexity from 9.62 to 7.77 and improving average zero-shot accuracy from 42.4% to 47.7%. Code is available this https URL.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.01153 [cs.LG] |
| (or arXiv:2610.01153v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01153 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Di He [view email]
[v1]
Thu, 1 Oct 2026 06:32:45 UTC (868 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org