Dataset Card GPT NL Public Corpus The GPT NL Public Corpus is the largest permissively licensed Dutch language resource available for large language model pretraining. It consists of 29 curated collections totaling over 524 billion tokens, including 36B Dutch, 207B English, 232B code, and 48B German/Danish tokens. All data is sourced under permissive licensing and redistributed under a CC BY license. For more details please refer to our Public Corpus article. Dataset Description Curated by: GPT NL Team Language(s) (NLP): Dutch, English, German, Danish, Frisian, Code License: CC BY. If you use (part of) the Public Corpus, please cite like stated below. Note : this is the license of the corpus. However, all datapoints have specific licenses attributed in the metadata field. These licenses are based on the license of the original source dataset that was curated or augmented for the purpose of creating the GPT NL public corpus. Author names are given for collections with a CC BY license. For Code datasets, we link the original Github link where the original license conditions can be found. Dataset Structure We have collected metadata of all 29 collections through doing interviews, doin…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy