FLEURS Dataset Description Fine Tuning script: pytorch/speech recognition Paper: FLEURS: Few shot Learning Evaluation of Universal Representations of Speech Total amount of disk used: ca. 350 GB Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven geographical areas: Western Europe : Asturian, Bosnian, Catalan, Croatian, Danish, Dutch, English, Finnish, French, Galician, German, Greek, Hungarian, Icelandic, Irish, Italian, Kabuverdianu, Luxembourgish, Maltese, Norwegian, Occitan, Portuguese, Spanish, Swedish, Welsh Eastern Europe : Armenian, Belarusian, Bulgarian, Czech, Estonian, Georgian, Latvian, Lithuanian, Macedonian, Polish, Romanian, Russian, Serbian, Slovak, Slovenian, Ukrainian Central Asia/Middle East/North Africa : Arabic, Azerbaijani, Hebrew, Kazakh, Kyrgyz, M…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy