TxT360: A Top Quality LLM Pre training Dataset Requires the Perfect Blend Changelog Version Details v1.1 Added new data sources: TxT360 BestOfWeb, TxT360 QA, europarl aligned, and wikipedia extended. Details of v1.1 Additions TxT360 BestOfWeb: This is a filtered version of the TxT360 dataset, created using the ProX document filtering model. The model is similar to the FineWeb Edu classifier, but also assigns an additional format… See the full description on the dataset page: https://huggingface.co/datasets/LLM360/TxT360.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy