💬 FineTranslations The world's knowledge in 1+1T tokens of parallel text Table of Contents 💬 FineTranslations What is it? What is it for? Languages and available subsets How to download and use 💬 FineTranslations + Using 🏭 datatrove + Using huggingface hub + Using datasets Dataset processing steps + 1. Sourcing the data + 2. Running translation at scale + 3. Post processing + 4. Edu filtering Dataset card for 💬 FineTranslations Dataset Description + Dataset Summary Dataset Structure + Data Instances + Data Fields + Data Splits Dataset Creation + Curation Rationale + Source Data + Data processing steps + Annotations + Personal and Sensitive Information and opt out Considerations for Using the Data + Social Impact of Dataset + Discussion of Biases + Other Known Limitations Additional Information + Licensing Information Citation Information What is it? This dataset contains over 1 trillion tokens of parallel text in English and 500+ languages. It was obtained by translating data from 🥂 FineWeb2 into English using Gemma3 27B. We relied on datatrove's inference runner to deploy a synthetic data pipeline at scale . Its checkpointing and VLLM lifecycle management features allowed us…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy