FineWeb2 HQ Dataset summary FineWeb2 HQ is a high quality, model filtered pretraining dataset derived as a subset of FineWeb2 , spanning 20 languages . It enables around 6x faster pretraining compared to the base dataset. FineWeb2 HQ was created by selecting the top 10% quality documents of FineWeb2 in each language, based on scores assigned by a deep learning classifier trained to identify structured and knowledge rich samples using XLM RoBERTa embeddings . Validation was performed by pretraining 1B parameter LLM models (llama like architecture) across multiple languages and writing systems (scripts). Evaluations on CMMLU (Chinese) and MMLU (German & French) demonstrate that FineWeb2 HQ matches FineWeb2 performance when trained with 6x fewer tokens, and outperforms it when fully trained . Additionally, improvements were observed across other benchmarks , such as outperforming its English cousins DCLM and FineWeb Edu. For more details, see our paper Enhancing Multilingual LLM Pretraining with Model Based Data Selection. Key features High quality selection : Top 10% of FineWeb2 documents by quality Multilingual coverage : 20 languages, ensuring diverse linguistic representation Mode…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy