跳到正文
arXiv:cs.LG· Shengye Tao, Yinzhu Cheng, Haihua Xie·· 2 天前AI 评分32

语言模型预训练中块绕过响应的持续深度排序研究

Persistent Depth Ordering amid Shifting Block-Bypass Responses in Language Model Pretraining

AI 导读

研究对五条已发布训练轨迹和11个模型-领域组合进行单块恒等绕过实验,发现块绕过响应在预训练过程中保持可识别的深度排序,但幅度会重新分布。在复现的Pythia训练中,局部缺失更新幅度增大而池化匹配下游响应减小,OLMo-2 7B则呈现不同平衡。结果表明单检查点干预响应需结合扰动路径随训练的演变来解读。

正文

View PDF HTML (experimental)

Abstract:Layer interventions are widely used to probe the internal organization of language models, yet most analyses examine a single training checkpoint even though model representations and computations evolve throughout pretraining. This leaves open which depth-dependent intervention responses reflect persistent organization and which are transient consequences of training. We study this question using single-block identity bypass on fixed teacher-forced contexts across five released trajectories and 11 model-domain combinations. We find that block-bypass responses retain recognizable depth ordering while their magnitudes redistribute: nearby checkpoints preserve stronger rank correspondence than distant ones, and large changes concentrate at positions that recur across text samples and transfer across evaluation domains. Controlled experiments further show that changes in the natural bypass effect cannot be reduced to a single downstream sensitivity: in replicated Pythia runs, local missing-update magnitude grows while the pooled matched downstream response decreases, whereas OLMo-2 7B exhibits a different balance. These matched responses also depend on perturbation strength and direction, without identifying targeted compensation. Together, our results show that longitudinal layer sensitivity is structured but not static, and that single-checkpoint intervention responses should be interpreted in the context of how the underlying perturbation pathway evolves during training.
Comments: 24 pages, 12 figures
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.01165 [cs.LG]
  (or arXiv:2610.01165v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.01165

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Rancho Tao [view email]
[v1] Thu, 1 Oct 2026 06:43:19 UTC (1,664 KB)

来源:arXiv:cs.LG · arxiv.org