FineWeb HQ Dataset Summary FineWeb HQ is a high quality, model filtered pretraining dataset derived as a subset of FineWeb . FineWeb HQ was created by selecting the top 10% of FineWeb documents based on a deep learning classifier trained to identify structured and knowledge rich samples . This classifier uses XLM RoBERTa embeddings to score documents. To validate our approach, we pretrained 1B parameter LLM models with a Llama like architecture across multiple languages and scripts. The results showed improvements on standard English benchmarks , with our dataset outperforming its English counterparts DCLM and FineWeb Edu. For its multilingual version, FineWeb2 HQ , evaluations on CMMLU (Chinese), MMLU (German), and MMLU (French) demonstrated that it matches FineWeb2's performance while trained on 6x fewer tokens and surpasses it when fully trained . Dataset Ours DCLM FW Edu FW : : : : : : : : : Average Rank 1.8333 2.3889 2.4444 3.3333 ARC (Challenge) 0.3550 0.3530 0.3850 0.3010 ARC (Easy) 0.6670 0.6470 0.6970 0.5880 CommonsenseQA 0.3870 0.4100 0.3770 0.3850 HellaSwag 0.6040 0.5960 0.5700 0.5930 MMLU 0.3400 0.3160 0.3470 0.3030 OpenBookQA 0.3860 0.3840 0.4180 0.3560 PIQA 0.7510 0.7…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy