TxT360: A Top Quality LLM Pre training Dataset Requires the Perfect Blend Changelog Version Details v1.1 Added new data sources: TxT360 BestOfWeb, TxT360 QA, europarl aligned, and wikipedia extended. Details of v1.1 Additions TxT360 BestOfWeb : This is a filtered version of the TxT360 dataset, created using the ProX document filtering model. The model is similar to the FineWeb Edu classifier, but also assigns an additional format score that considers how a document is formatted. TxT360 QA : Synthetic QA pairs generated for each document using Mistral 7B Instruct v0.3. QA pairs are appended to the end of every document in the format: The number of QA pairs may differ for each document, providing diverse question answering supervision. europarl aligned : Europarl v7 data processed to align English source text with parallel corpora in multiple languages. Each sample concatenates the same content in different languages. Steps include reading English source text, matching with parallel corpus data, and concatenating multilingual content for robust cross lingual training without any order., e.g.: wikipedia extended : An enhanced version of Wikipedia data that: Appends abstracts of outgoi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy