🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer Table of Contents 🍷 FineWeb What is it? What is being released? Changelog How to download and use 🍷 FineWeb + Using 🏭 datatrove + Using huggingface hub + Using datasets Breakdown by dump/crawl Dataset performance evaluation and ablations + Hyper parameters for ablation models + Ablation evaluation benchmarks + Comparison with other datasets Dataset card for 🍷 FineWeb Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information + Future work + Citation Information What is it? The 🍷 FineWeb dataset consists of more than 18.5T tokens (originally 15T tokens) of cleaned and deduplicated english web data from CommonCrawl. The data processing pipeline is optimized for LLM performance and ran on the 🏭 datatrove library, our large scale data processing library. 🍷 FineWeb was originally meant to be a fully open repli…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy