跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yuchen Li, Mingyu Du, Zongqi Fan, Ken-Tye Yong, Nguyen H. Tran·· 14 小时前AI 评分29

预训练模型为何出现训练-验证分离?更新压力密度动力学给出解释

Why Does Train-Validation Separation Emerge? Update-Pressure Density Dynamics in Pretrained Backbones

AI 导读

研究提出"更新压力密度"动力学解释预训练模型微调中的训练-验证分离:持续拟合会把更新需求从可广泛复用的支撑转向迁移性更弱的窄支撑。在 ResMLP 构造层级中,样本私有特征占比从 p=.3 升至 .5、.7 时,最终平均准确率差距由 .185 升至 .331、.527;RoBERTa、DeBERTa、Qwen 在六个数据集共 90 次运行中,训练探针加权的类内与整体离散度读数均与准确率差距正相关。

正文

View PDF HTML (experimental)

Abstract:Train-validation separation is the evolving difference between performance on observed training examples and a finite held-out validation set. We propose a dynamic structural account of how this gap develops during adaptation of pretrained models: continued fitting can shift update demand from broadly reusable support toward narrower support with weaker held-out transfer. A conditional local model links this shift to increasing heterogeneity in gradient allocation and train-validation separation. Fixed training probes make this structural evolution observable without validation examples entering the readouts; held-out performance is used separately to evaluate its relation to the gap. In a constructed hierarchy implemented with a residual multilayer perceptron (ResMLP), increasing the target share of example-private features from $p=.3$ to $.5$ to $.7$, while preserving the relative mixture $1{:}2{:}3{:}4$ among the four shared feature levels, increases the final mean accuracy gap from $.185$ to $.331$ to $.527$ across five runs per condition. Masked-input losses measured separately on training and validation examples expose the corresponding transfer asymmetry. The natural language processing (NLP) analysis uses 10-epoch runs of RoBERTa, DeBERTa, and Qwen on six datasets (90 runs): the training-probe-weighted within-class and overall dispersion readouts each have positive raw and smoothed level correlations with the accuracy gap in all 90 runs. Raw changes paired at approximately one-epoch intervals remain positively associated in 86/90 and 87/90 runs, respectively. A 40-epoch ResNet-18 study tests both readouts on three vision datasets. Together, controlled simulation, NLP, and vision support the dynamic structural account across settings, with real-model evidence testing its observable predictions under the specified monitors.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.01425 [cs.LG]
  (or arXiv:2610.01425v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01425

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yuchen Li [view email]
[v1] Thu, 1 Oct 2026 10:25:57 UTC (2,443 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org