๐ฅ FineWeb2 A sparkling update with 1000s of languages Table of Contents ๐ฅ FineWeb2 What is it? Languages and available subsets + How many tokens? Changelog How to download and use ๐ฅ FineWeb2 + Using ๐ญ datatrove + Using huggingface hub + Using datasets Dataset processing steps + Language Identification ๐ + Deduplication ๐๏ธ + Data Filtering ๐งน + PII Anonymization and fixes ๐ญ Dataset performance evaluation and ablations + Hyper parameters for ablation models + Score normalization + Global scores across languages + Comparison with other datasets Dataset card for ๐ฅ FineWeb2 Dataset Description + Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information and opt out Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information + Future work and community initiatives Citation Information What is it? This is the second iteration of the popular ๐ท FineWeb dataset, bringing high quality pretraining data to over 1000 ๐ฃ๏ธ languages. The ๐ฅ FineWeb2โฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy