INFINI NEWS Corpus A multilingual news corpus extracted from Common Crawl CC News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers Aug 2016 – Apr 2026. At a glance Articles 1 357 027 742 Distinct hostnames 133 565 Text size (uncompressed) 3.4 TB Estimated tokens ~1.19 trillion (cl100k base) Languages — GlotLID (ISO 639 3) 1 172 Languages — CommonLingua (334 class) 331 Files 48 978 parquet shards, partitioned year=YYYY/month=MM On disk size ~1.7 TB (zstd parquet) License CC BY 4.0 Articles per year year articles months : 2016 7.4 M Aug–Dec 2017 71.6 M full 2018 92.0 M full 2019 129.8 M full 2020 185.2 M full 2021 196.7 M full 2022 201.9 M (peak) full 2023 182.1 M full 2024 143.1 M full 2025 120.3 M full 2026 26.8 M Jan–Apr Top languages Detected with two independent identifiers — GlotLID (Kargaran et al. 2023) and CommonLingua (Pleias / GSMA 2026); counts below are from CommonLingua (334 class, byte level). ISO 639 3 articles share : : eng 521.3 M 38.4 % spa 135.2 M 9.96 % rus 88.8 M 6.55 % deu 86.7 M 6.39 % ita 70.8 M 5.21 % fra 55.3 M 4.07 %…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy