π CS FLEURS: A Massively Multilingual and Code Switched Speech Dataset π Overview CS FLEURS is a new dataset for developing and evaluating code switched speech recognition and translation systems beyond high resourced languages. 113 unique code switched language pairs across 52 languages 300 hours of speech data, both read and synthetic π Dataset Statistics CS FLEURS consists of the following subsets: Read Test: 14 X English pairs, read speech XTTS Train: 16 X English pairs, generative TTS XTTS Test1: 16 X English pairs, generative TTS XTTS Test2: 60 {Arabic, Chinese, Hindi, Spanish} X pairs, generative TTS MMS Test: 45 X English pairs, concatenative TTS Statistic Read Test XTTS Train XTTS Test1 XTTS Test2 MMS Test Duration (hours) 17 128 36 42 56 Tokens (words) 128k 889k 257k 300k 315k Matrix Langs 14 16 16 4 45 Embedded Langs 1 1 1 15 1 Total CS Pairs 14 16 16 60 45 Same Script Pairs 7 10 10 10 22 Distinct Script Pairs 7 6 6 51 23 π½ How to Download Clone with Git LFS: Or load directly in π€ Datasets: π License This dataset is licensed under CC BY 4.0 for Non Commercial use. π Citation If you use this dataset, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy