跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yimiao Yu, Florentin Guth·· 14 小时前AI 评分34

迁移学习研究:定位记忆任务间的迁移效应

Localizing Transfer Between Memorization Tasks

AI 导读

研究者在随机输入-输出映射的记忆任务间发现两种迁移模式:等效迁移,即每增加一个预训练 epoch 约节省一个下游微调 epoch;以及非等效迁移,即在不匹配任务上预训练甚至比直接训练下游任务更高效。通过消融实验,迁移被分解为最后一层由幅度驱动的“平凡”迁移和部分归因于其他层协方差的“非平凡”结构驱动迁移。该成果已被 NeurIPS 2026 Workshop 接收为 poster。

正文

View PDF HTML (experimental)

Abstract:A central puzzle in transfer learning is why pre-training on one task can accelerate training or improve performance on another task, and what mechanisms underlie this transfer. In this work, we examine the transfer between memorization tasks of random input-output mappings. We find two surprising transfer patterns: equivalent transfer, where each additional pre-training epoch saves approximately one downstream fine-tuning epoch; and non-equivalent transfer, where pre-training on a mismatched task can be even more efficient than directly training on the downstream task itself. Through ablation experiments, we decompose and localize the transfer into two separate effects: a "trivial" magnitude-driven transfer in the last layer, and a "non-trivial" structure-driven transfer, partially attributable to the covariance of the other layers. These results advance our understanding of the underlying mechanisms of transfer learning and have the potential to lead to principled pre-training strategies.
Comments: 17 pages, 10 figures. Accepted as a poster at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.00771 [cs.LG]
  (or arXiv:2610.00771v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00771

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yimiao Yu [view email]
[v1] Wed, 30 Sep 2026 22:05:38 UTC (3,487 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org