跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Peter Nutter, Dani Roytburg, Cl\'ement Dumas, Jinghua Ou, Shi Feng·· 14 小时前AI 评分44

预训练干预新方法 Grafting:跨 checkpoint 移植模型信念

Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints

AI 导读

研究者提出 Grafting 方法,在预训练 checkpoint 上训练 SDF 适配器,再将学到的权重更新加到已后训练模型上,从而近似忠实的预训练干预并复用现有后训练。

正文

View PDF HTML (experimental)

Abstract:Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideally, synthetic documents would be mixed into pre- or mid-training, but every change to a pre-training corpus must be followed by a full post-training run before its effect can be measured, making iteration slow and expensive. Common practice instead applies SDF to an already post-trained model. This is known to leave artifacts and degrade capabilities, and, as we show, it makes the model treat fabricated entities unrelated to the documents as real, a failure we call reality drift. We propose grafting: train the SDF adapter on the pre-trained checkpoint, then add the learned weight update to the post-trained model, which approximates the faithful approach while reusing the existing post-training. We demonstrate this by installing false facts, training misaligned model organisms and applying a constitutional mid-training intervention, across model families up to 284B parameters. Grafting installs the target belief as strongly as SDF on the post-trained model while reducing both reality drift and the loss of preference coherence by more than half on average, and it stays closer to a faithful mid-training run. Because grafting requires no post-training, the same adapter can be applied to any later checkpoint, enabling researchers to iterate quickly on pre-training interventions at the cost of a single fine-tuning run.
Comments: 78 pages. Code: this https URL
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.00767 [cs.LG]
  (or arXiv:2610.00767v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.00767

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Peter Nutter [view email]
[v1] Wed, 30 Sep 2026 22:02:46 UTC (2,626 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org