C4 Dataset Description Paper: https://arxiv.org/abs/1910.10683 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en , en.noclean , en.noblocklist , realnewslike , and multilingual (mC4). For reference, these are the sizes of the variants: en : 305GB en.noclean : 2.3TB en.noblocklist : 380GB realnewslike : 15GB multilingual (mC4): 9.7TB (108 subsets, one per language) The en.noblocklist variant is exactly the same as the en variant, except we turned off the so called "badwords filter", which removes all documents that contain words from the lists at https://github.com/LDNOOBW/List of Dirty Naughty Obscene and Otherwise Bad Words. How do I download this? Using 🤗 Datasets Since this dataset is big, it is encouraged to load it in streaming mode using streaming=True , for example: You can also load and mix multiple languages: Using Dask Using Git This will download 13TB to your local drive. If you want to be more precise with what you are downloading, follow these commands instead: The git clone command in th…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy