arXiv:cs.CL· Thomas Gebhart, Russell J. Funk·· 3 小时前
SciTBERT:面向科技语言处理的时间一致语言模型系列
SciTBERT: A family of chronologically consistent language models for scientific and technological language processing
AI 导读
SciTBERT 是一系列时间一致的 BERT 衍生语言模型,训练数据来自科学论文、专利和高质量教育网页文本,数据截止日期覆盖 2013 至 2025 年的每一年。
正文
Abstract:Pre-trained transformer models are increasingly being used to study scientific and technological progress. Encoders tuned to paper or patent text outperform general-purpose models on downstream classification, regression, and proximity tasks within science and technology. However, the applicability of these models for studying time-dependent or archival properties of science, technology, and their interface is limited due to lookahead and domain biases inherent to these pre-trained models. These limitations arise from training on corpora with unconstrained chronological and text source distributions. We introduce SciTBERT: a family of chronologically consistent BERT-derived language models trained on text from scientific papers, patents, and high-quality educational web text with training data cutoff dates spanning each year between 2013 and 2025. We also post-train these models in a chronologically-consistent manner using paper and patent citations, creating SciTBERT-CI model family. We find that these models generally outperform predecessor domain-specific encoder models even when training data is limited by early year restrictions in the corpus. To further investigate the extent to which this class of models can learn representations that bridge science and technology, we introduce the PatRepEval benchmark, a suite of patent-related text embedding tasks at the science-technology interface. Performance in a variety of classification, regression, and retrieval tasks spanning papers and patents highlights the importance of aligning encoder model representations with the domain distributions of their downstream tasks, and chronologically consistent encoders can match or exceed models trained without temporal constraints.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12207 [cs.CL] |
| (or arXiv:2610.12207v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12207 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Thomas Gebhart [view email]
[v1]
Thu, 8 Oct 2026 15:57:45 UTC (176 KB)
来源:arXiv:cs.CL · arxiv.org