Helsinki-NLP/fineweb-edu-translated
fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are based on OPUS-MT and HPLT-MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space-searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages.
In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb-edu data set with its 350B token release.
More information about how the data has been produced can be found on https://github.com/Helsinki-NLP/translate-fineweb. The corpus is also available with aligned sentences in synOPUS: https://opus.nlpl.eu/synthetic/transweb-edu.php
Supported Languages
| Langid | Language |
|---|---|
| bos | Bosnian |
| bul | Bulgarian |
| cat | Catalan |
| ces | Czech |
| dan | Danish |
| deu | German |
| ell | Modern Greek |
| eng | English |
| est | Estonian |
| eus | Basque |
| fin | Finnish |
| fra | French |
| gle | Irish |
| glg | Galician |
| hrv | Croatian |
| hun | Hungarian |
| isl | Icelandic |
| ita | Italian |
| kat | Georgian |
| lav | Latvian |
| lit | Lithuanian |
| mkd | Macedonian |
| mlt | Maltese |
| nld | Dutch |
| nno | Norwegian Nynorsk |
| nob | Norwegian Bokmål |
| pol | Polish |
| por | Portuguese |
| ron | Romanian |
| slk | Slovak |
| slv | Slovenian |
| spa | Spanish |
| sqi | Albanian |
| srp_Cyrl | Serbian (cyrillic script) |
| swe | Swedish |
| tur | Turkish |
| ukr | Ukrainian |
Citation Information
Please acknowledge the source when using the data and, please, cite the following article if you use any part of this corpus in your own work:
@article{tiedemann2023democratizing,
title={Democratizing neural machine translation with {OPUS-MT}},
author={Tiedemann, J{\"o}rg and Aulamo, Mikko and Bakshandaeva, Daria and Boggia, Michele and Gr{\"o}nroos, Stig-Arne and Nieminen, Tommi and Raganato, Alessandro and Scherrer, Yves and Vazquez, Raul and Virpioja, Sami},
journal={Language Resources and Evaluation},
number={58},
pages={713--755},
year={2023},
publisher={Springer Nature},
issn={1574-0218},
doi={10.1007/s10579-023-09704-w}
}
Translation Models
The following translation models have been used for creating the data:
Acknowledgements
None of this would be possible without the enormous work done by Common Crawl providing the essential data that most open datasets for language modeling are based on. Furthermore, we are also grateful for the data preparation work done by Hugging Face and the community on top of the crawled data from Common Crawl published under the label of fineweb-edu. Important for this work is also the availability of parallel data through OPUS and the public translation models based on that data. Furthermore, this project was supported by the European Union's Horizon Europe research and innovation programme through the HPLT project under grant agreement No 101070350. Finally, we also want to acknowledge the computational resource made available from the Finnish national allocation for the LUMI supercomputer (https://www.lumi-supercomputer.eu) through the extreme scale project MaMuLaM: Massively Multilingual Language Models. None of the translated data would exist without this infrastructure and the compute.