arXiv:cs.CL· Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer·· 4 小时前AI 评分47
MoE 模型对重复数据比稠密模型更易过拟合
Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
AI 导读
研究发现,在 80M 至 1B 激活参数(8.5B 总参数)规模下,Mixture-of-Experts(MoE)模型在数据重复训练时退化速度快于稠密模型,且稀疏度越高越严重。80M 稠密模型可重复数据超 8x 而几乎无退化,MoE 在 4x 时即开始受损,32x 后性能低于稠密模型。强掩码正则化可使 MoE 在数据重复超 64 次时仍优于稠密模型,但无任何方法能完全媲美全唯一数据训练。
正文
Abstract:As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8x with minimal degradation, MoEs instead begin to suffer at 4x, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to underperform dense models after 32x. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL) |
| Cite as: | arXiv:2609.11917 [cs.LG] |
| (or arXiv:2609.11917v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.11917 arXiv-issued DOI via DataCite |
Submission history
From: Margaret Li [view email]
[v1]
Thu, 10 Sep 2026 17:57:33 UTC (1,429 KB)
[v2]
Wed, 7 Oct 2026 17:52:42 UTC (1,408 KB)
来源:arXiv:cs.CL · arxiv.org