π FinePDFs Edu 350B+ of highly educational tokens from PDFs π What is it? π FinePDFs Edu dataset consists of 350B+ tokens of educational PDFs filtered from π FinePDFs dataset covering 69 languages. FinePDFs was created using the formula inspired from FineWeb Edu, we developed an educational quality classifier using annotations generated by Qwen3 235B A22B Instruct 2507 for each of 69 languages present in this dataset. We then used this classifier to retain only the most educational web pages. FinePDFs Edu outperforms FinePDFs on popular benchmarks and shows the power of classifiers trained on synthetic data. The Dataset Curation section details the process for creating the dataset. While it might seem that the dataset is an order of magnitude smaller than FineWeb Edu, unlike its web ancestor, this dataset is globally deduplicated! What is being released? Along with the dataset, which includes all filtered CommonCrawl dumps since CC MAIN 2013 20 to CC MAIN 2025 08 , we also release: The educational classifier used for the filtering (for each language) The dataset with educational (and 3 other) labels by Qwen3 235B A22B Instruct 2507 for English. The dataset with educational labeβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy