FineWeb Edu 100BT (Shuffled) A globally shuffled version of HuggingFaceFW/fineweb edu 100BT. Part of the Smol Data collection — tried and tested mixes for strong pretraining. Dataset Description This dataset contains the same ~100B tokens as fineweb edu 100BT but with all documents globally shuffled (seed=42). Use this version when you need randomized document ordering for pretraining. How It Was Created The unshuffled dataset was loaded into memory, shuffled with dataset.shuffle(seed=42) , and re uploaded with 100 shards. See the smol data.py script for details. Usage Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy