arXiv:cs.AI· Xiaoxi Luo, Zachary Shinnick, Niclas Griesshaber, Yixuan Wang, Junchi Yu, Freda Shi, Philip Torr, Yao Lu·· 4 小时前AI 评分43
TypewriterLM:在 1913 年前历史文本上预训练的 7.24B 语言模型
A Language Model from 1913: Pretraining on Historical Text
AI 导读
研究者预训练了 TypewriterLM,一个 7.24B 参数、知识截止于 1913 年的语言模型,证明在数据受限条件下用历史文本预训练仍能保持合理的语言理解能力。
正文
Abstract:While modern language models increasingly rely on ever-larger web corpora, we show that pretraining on historical text (e.g., pre-1913 text) in a data-constrained setting can produce a temporally grounded language model that still shows reasonable performance on language understanding. However, developing History LMs requires addressing challenges in data quality, preventing temporal leakage in post-training, and constructing temporally aligned evaluations. We address these challenges and pretrain TypewriterLM, a 7.24B-parameter model with a 1913 knowledge cutoff. We construct TypewriterCorpus, a 54B-token historical corpus with extensive temporal filtering, propose lexically grounded instruction tuning that constrains all responses to vocabulary from historical source documents, and introduce History-Event, a benchmark of 2,344 events for evaluating both competence and cutoff adherence. We release TypewriterLM and all associated resources to support future research on History LMs.
| Comments: | Accepted by EMNLP 2026 |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.02991 [cs.CL] |
| (or arXiv:2606.02991v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.02991 arXiv-issued DOI via DataCite |
Submission history
From: Xiaoxi Luo [view email]
[v1]
Tue, 2 Jun 2026 00:59:06 UTC (95 KB)
[v2]
Fri, 2 Oct 2026 04:21:31 UTC (113 KB)
来源:arXiv:cs.AI · arxiv.org