跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Zachary Shinnick, Hemanth Saratchandran, Damien Teney, Anton van den Hengel·· 5 小时前AI 评分43

抽象数据预训练的影响可超越困惑度指标

Lasting Effects of Abstract Pretraining Beyond Perplexity

AI 导读

小型语言模型预训练时,仅用1%的token进行抽象栈操作任务热身,可将MUSIQUE多跳问答提升最多3.9个F1点,HOTPOTQA和2WIKIMULTIHOPQA也有增益,但语言建模困惑度相当。该增益具有特异性:仅改善顺序推理链,对独立事实的组合或比较任务无一致收益;且时机关键,预训练后接触抽象数据则完全无效。

正文

View PDF HTML (experimental)

Abstract:Language models are typically pretrained from random initialization. Recent work challenges this convention, showing that a brief warm-up on abstract, algorithmically generated data can provide a better starting point for subsequent learning of natural language. In this paper, we show that in small language models, such a warm-up improves specific capabilities that are not reflected in language-modeling perplexity. Our warm-up uses an abstract stack-manipulation task that requires compositional and state-tracking capabilities. Allocating as little as 1% of pretraining tokens to this data improves multi-hop question answering by up to 3.9 F1 points on MUSIQUE, with additional gains on HOTPOTQA and 2WIKIMULTIHOPQA despite comparable language-modeling perplexity. Controlled experiments show that the warm-up substantially accelerates the acquisition of deeper reasoning chains. We also explore what drives this transfer. First, the structure of the data matters: replacing the stack task with a queue fails to produce the same gains. Second, the gains are specific: performance improves on sequential reasoning chains, with no consistent benefit on tasks that combine or compare independent facts. Third, timing matters: mixing abstract data with natural language is far less effective than an initial dedicated phase, and exposure after pretraining completely removes the benefits. The early advantage persists through billions of subsequent language tokens. These results show that early abstract training can reliably shape the capabilities language models later acquire.
Comments: Project page: this https URL
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2609.38764 [cs.LG]
  (or arXiv:2609.38764v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.38764

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zachary Shinnick [view email]
[v1] Wed, 30 Sep 2026 01:48:40 UTC (1,345 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org