跳到正文
arXiv:cs.CL· Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Ir\`ene Girard, David Mach, Anastasia Stasenko, Ivan P. Yamshchikov·· 6 小时前AI 评分60

Common Corpus:面向 LLM 预训练的最大开放数据集论文

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

AI 导读

论文介绍 Common Corpus,一个约两万亿 token 的开放 LLM 预训练数据集,数据均为无版权或开放许可内容,涵盖多语言(含低资源语言)和大量代码数据。作者在数据集上训练两个小模型,性能与同规模其他模型相当,并公开了数据来源、筛选与整理细节,论文被 ICLR 2026 接收为 Oral。

正文

View PDF HTML (experimental)

Abstract:Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in pre-training datasets. In addition, it includes a large amount of code data. The diversity of data sources in terms of covered domains and time periods opens up the paths for both research and entrepreneurial needs across diverse areas of knowledge. In this paper, we present the detailed provenance of data assembling and the details of dataset filtering and curation. We train two small language models on Common Corpus and find that they perform comparably to other models of their size, indicating that our dataset is suitable for multilingual pretraining. Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2506.01732 [cs.CL]
  (or arXiv:2506.01732v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2506.01732

arXiv-issued DOI via DataCite

Journal reference: ICLR 2026 (Oral)

Submission history

From: Pavel Chizhov [view email]
[v1] Mon, 2 Jun 2025 14:43:15 UTC (695 KB)
[v2] Mon, 20 Apr 2026 13:44:50 UTC (828 KB)
[v3] Fri, 15 May 2026 14:33:53 UTC (828 KB)
[v4] Mon, 5 Oct 2026 18:35:02 UTC (828 KB)

来源:arXiv:cs.CL · arxiv.org