Helsinki NLP/nemotron cc translated nemotron cc tanslated is a collection of automatically translated documents from nemotron cc taken out of the high quality subset. Translations are based on OPUS MT and HPLT MT models. The data in v1.0 covers 156,431,999 documents with over 70 billion space searated tokens of English data translated into 36 languages. The total v1.0 data set includes over 2.4 trillion tokens and the translated documents are aligned across all languages. v1.1 includes a second batch of nemotron cc hq data (94 billion space separated tokens) translated into 9 languages: Bulgarian (bul), Czech (ces), Estonian (est), Finnish (fin), Irish (gle), Romanian (ron), Swedish (swe), Turkish (tur) and Ukrainian (ukr). The translation models used for the translation are the same as for v1.0 (listed below) but translation has been done on CPU using ctranslate2 and beam search of size 1 instead of using the original Marian NMT models on GPU with beam search of size 4. The additional files in v1.1 are marked with run2 in their file names. Version 1 is marked with run1 . More information about how the data has been produced can be found on https://github.com/Helsinki NLP/translate…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy