跳到正文
arXiv:cs.LG· Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan·· 3 小时前AI 评分42

基于 Fisher 引导的子模数据选择:面向大语言模型持续预训练

Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

AI 导读

研究者提出一种 Fisher 感知的持续预训练数据选择器,将候选样本梯度分解为锚点分量与前沿分量,并用对数行列式子模目标在单次流式流程中优化。在 TinyLlama-1.1B 和 Llama-3.1-8B 的医疗数据 CPT 上,该方法在提升目标域质量的同时限制了预训练基准上的遗忘。

正文

View PDF HTML (experimental)

Abstract:Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual pre-training (CPT), it becomes a forgetting-control problem: a poorly chosen target-domain corpus can overwrite capabilities encoded in the pretrained checkpoint. Existing CPT practice either scores candidates with parameter-agnostic scalars such as perplexity, or mitigates forgetting by spending many extra general-domain replay tokens. Neither strategy directly asks how training on a candidate will move the model parameters. We show that loss-based selection causes the post-CPT Fisher diagonal to drift downward on exactly the high-Fisher coordinates the pretrained model had committed to, while leaving low-Fisher coordinates largely untouched. This asymmetry exposes a parameter-space mechanism for catastrophic forgetting. Motivated by this observation, we propose a Fisher-aware CPT selector that decomposes each candidate's gradient into an anchor component, which measures perturbation along committed parameter directions, and a frontier component, which measures update capacity in unconstrained low-Fisher subspaces. We aggregate these signals with a log-determinant submodular objective and optimize it in a single pass using a scalable streaming data selection pipeline. On TinyLlama-1.1B and Llama-3.1-8B CPT over medical data, our selector improves target-domain quality while bounding forgetting on held-out pretraining benchmarks. Most importantly, it is substantially more token-efficient than forgetting-aware replay. 1B selected tokens already outperform the replay strategy trained with 10B tokens on both adaptation and forgetting, giving a 10x token-efficiency advantage.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.02593 [cs.LG]
  (or arXiv:2610.02593v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.02593

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhenghao Zhao [view email]
[v1] Thu, 1 Oct 2026 23:44:09 UTC (323 KB)

来源:arXiv:cs.LG · arxiv.org