arXiv:cs.CL· Mingzi Cao, Xi Wang, Nikolaos Aletras·· 3 小时前
OT-DUS:用最优传输做 LLM 深度扩容
Optimal Transport Depth Up-Scaling
AI 导读
研究者提出 Optimal Transport Depth Up-Scaling(OT-DUS),利用最优传输理论在相邻基座层间逐模块对齐并融合功能对应的神经元,以构建新层,缓解现有深度扩容方法因复制或平均基座层导致的神经元置换错配问题。在持续预训练和监督微调下,OT-DUS 在不同模型规模和模型族上于通用与专业领域均优于现有方法。分析还发现,将新层插入更高位置可同时带来更强性能与更高训练效率。
正文
Abstract:Pre-training Large Language Models (LLMs) from scratch at larger scales yields remarkable performance but incurs substantially high training costs. Depth up-scaling provides an efficient alternative by inserting new layers into a pre-trained LLM, avoiding training from scratch. However, most existing methods copying or averaging base layers for new layer, which misalign functionally corresponding neurons, leading to neuron permutation mismatch that harms performance. To address this issue, we propose Optimal Transport Depth Up-Scaling (OT-DUS), which leverages Optimal Transport (OT) theory to align and fuse functionally corresponding neurons module by module in adjacent base layers for new layer construction. OT-DUS achieves better overall performance in both general and specialized domains than existing methods for continual pre-training and supervised fine-tuning across different model sizes and model families. Our analysis of insertion strategies further finds that inserting new layers at higher positions yields not only stronger performance but also improved training efficiency.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2508.08011 [cs.CL] |
| (or arXiv:2508.08011v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2508.08011 arXiv-issued DOI via DataCite |
Submission history
From: Mingzi Cao [view email]
[v1]
Mon, 11 Aug 2025 14:15:33 UTC (336 KB)
[v2]
Wed, 7 Oct 2026 23:52:52 UTC (342 KB)
来源:arXiv:cs.CL · arxiv.org