Common Pile v0.1 — Parquet Consolidated Description This dataset bundles all “raw” corpora from the Common Pile v0.1 Raw Data collection , converted to Apache Parquet and consolidated in a single repository. Nothing has been filtered or modified; the only changes are: Format: original JSON → Parquet Layout: many repositories → one consolidated dataset Extra column: a len category bucket for quick length based filtering Only the three original columns ( id , text , source ) are carried over; len category is derived from text length. Dataset Schema Column Type Notes id string Original document ID text string UTF 8 plain text source string Name of the originating corpus (e.g. library of congress ) len category string Bucketed document length (bytes) License Issues Licensing follows the individual corpora. While we aim to produce datasets with completely accurate licensing information, license laundering and inaccurate metadata can cause us to erroneously assign the incorrect license to some documents (for further discussion of this limitation, please see our paper). If you believe you have found an instance of incorrect licensing in this dataset, please start a discussion on this repo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy