Common Corpus Full paper ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open : contains only data that is either uncopyrighted or freely licensed Traceable : each individual document is associated with documented contextual information, including licensed use or lack of copyright. Multilingual : mostly representing English and French data, but contains data for 8 languages with more than 10 billion tokens (German, Spanish, Italian, Polish, Greek, Latin) and 33 languages with more than 1 billion tokens. Diverse : consisting of scientific articles, government and legal documents, code, and cultural heritage data, including books and newspapers Extensively Curated : spelling and formatting has been corrected from digitized texts, harmful and toxic content has been removed, and content with low educational content has also been re…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy