π FineWeb Edu score 2 1.3 trillion tokens of the finest educational data the π web has to offer What is it? π FineWeb Edu dataset consists of 1.3T tokens (FineWeb Edu) and 5.4T tokens of educational web pages filtered from π· FineWeb dataset. This is the 5.4 trillion version. Note: this version uses a lower educational score threshold = 2, which results in more documents, but lower quality compared to the 1.3T version. For more details check the FineWeb blog post. To enhance FineWeb's quality, we developed an educational quality classifier using annotations generated by LLama3 70B Instruct. We then used this classifier to retain only the most educational web pages. FineWeb Edu outperforms FineWeb on popular benchmarks and shows the power of classifiers trained on synthetic data. The Dataset Curation section details the process for creating the dataset. What is being released? Along with the dataset, which includes all filtered CommonCrawl dumps since 2013, we also release the educational classifier used for the filtering as well as the code for training it and running inference at: https://github.com/huggingface/cosmopedia/tree/main/classification. Changelog Previous versions reβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy