NB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large scale collection of web crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre print. The Cleaned variant of HPLT Datasets v2.0 This is the variant of the HPLT Datasets v2.0 converted to the Parquet format semi automatically when being uploaded here. The original JSONL files (which take ~4x fewer disk space than this HF version) and the larger non cleaned version can be found at https://hplt project.org/datasets/v2.0. Dataset Performance Internal Evaluation We conducted the FineWeb style ablation studies within the HPLT project with the focus on one high resource and one low resource language: English and Norwegian. We train 1.7B decoder only LMs using 100B/30B tokens sampled from the English/Norwegian parts of our HPLT v2 dataset respectively. We replicate the FineWeb corpora comparison design and train the models with a fixed pretraining se…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy