跳到正文
arXiv:cs.LG· Shih-Cheng Huang, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun·· 4 小时前AI 评分36

αTransfer:用小模型迁移合并系数实现高效模型合并

$\alpha$Transfer: Coefficient Transfer for Efficient Model Merging

AI 导读

αTransfer 通过在小型代理模型上搜索最优合并系数并直接迁移到更大目标模型,实现高效模型合并。实验显示,该方法在视觉 Transformer 上提速 6 倍、内存降低 70%,在大语言模型上提速 20 倍、内存降低 85%,同时保持性能相当。该范式已在多种合并方法、模型族和任务上得到验证,代码尚未开源。

正文

View PDF HTML (experimental)

Abstract:Model merging offers a promising solution for combining multiple fine-tuned checkpoints into a single model through parameter arithmetic. However, finding optimal merging coefficients requires an extensive search that becomes prohibitively expensive as models scale in both size and number, due to high memory requirements and combinatorial growth in the search space. We show that, within the same model family, models exhibit highly congruent performance distributions over merging coefficients across different model sizes. This distributional similarity enables a practical paradigm we call \textit{$\alpha$Transfer}: searching for optimal coefficients on a small proxy model, then directly transfer them to larger target models. We verify $\alpha$Transfer across multiple merging methods, model families, and tasks. Experimental results demonstrate a 6$\times$ speedup and 70\% memory reduction on vision transformers, and a 20$\times$ speedup and 85\% memory reduction on large language models, while maintaining comparable performance. Our findings establish $\alpha$Transfer as an efficient and generalizable approach to scaling model merging.
Comments: Under review
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2610.07819 [cs.LG]
  (or arXiv:2610.07819v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.07819

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Shih-Cheng Huang [view email]
[v1] Tue, 6 Oct 2026 06:13:01 UTC (172 KB)

来源:arXiv:cs.LG · arxiv.org