Helsinki NLP/fineweb edu translated fineweb edu tanslated is a collection of automatically translated documents from fineweb edu. Translations are based on OPUS MT and HPLT MT models. The data in v1.0 covers 36,704,000 documents with over 28 billion space searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 960 billion tokens and the translated documents are aligned across all languages. In the v1.1 release, additional translations are added for Czech (ces), Ukrainian (ukr) and Finnish (fin). For Czech and Ukrainian, this release doubles the data and for Finnish, we include translations for the entire fineweb edu data set with its 350B token release. More information about how the data has been produced can be found on https://github.com/Helsinki NLP/translate fineweb. The corpus is also available with aligned sentences in synOPUS: https://opus.nlpl.eu/synthetic/transweb edu.php Supported Languages Langid Language bos Bosnian bul Bulgarian cat Catalan ces Czech dan Danish deu German ell Modern Greek eng English est Estonian eus Basque fin Finnish fra French gle Irish glg Galician hrv Croatian hun Hungarian isl Icelandic ita Italian kat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy