EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large scale multilingual speech corpus containing high quality aligned parliamentary speech across 22 European languages. The dataset was constructed by processing parliamentary proceedings using a robust alignment pipeline that handles diverse audio formats and non verbatim transcripts. More information can be found in the paper. Dataset Summary Languages : 22 European languages (see detailed breakdown below) Total aligned hours : ~78,100 hours of initially aligned speech text data Quality filtered subsets : CER < 30%: approximately 61,000 hours CER < 20%: approximately 50,500 hours (this is the primary subset provided directly through the Hugging Face Datasets interface for all languages) CER < 10%: approximately 32,200 hours Domain : Parliamentary proceedings (formal speaking style) Audio segment length : Typically 3 20 seconds Format : Audio segments with paired transcriptions Languages EuroSpeech provides substantial data for previously under resourced languages: 19 languages exceed 1,000 hours of data (CER < 20%) 22 languages exceed 500 hours of data (CER < 20%) Language Code Total Aligned (h) CER < 30\% (h) CER <…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy