WorldSpeech A multilingual ASR dataset containing over 65k hours of human transcribed speech across 127 language region variants, drawn from national parliaments, public broadcasters, public domain audiobooks, and international institutions. Rows consist of 24 kHz speech utterances paired with a human provided transcript, an aligned ASR transcript, character error rate (CER) between the two, a WADA SNR estimate, and four DNSMOS P.835 quality scores. Dataset Overview Metric Value : : Total hours 65,072 Distinct languages 88 (76 above 10 hours) Language variations 127 Languages ≥ 1,000 h 24 Languages ≥ 500 h 28 Languages ≥ 200 h 37 Languages ≥ 50 h 53 Avg DNSMOS P.835 OVR 2.83 Sample rate 24 kHz How to use Streaming Schema Field Description audio OGG Opus decoded to {"array": np.float32[N], "sampling rate": 24000} human transcript Human provided transcript asr transcript ASR output used for alignment and CER computation cer Character error rate between asr transcript and human transcript snr WADA SNR estimate, dB dnsmos sig DNSMOS P.835 signal quality dnsmos bak DNSMOS P.835 background noise dnsmos ovr DNSMOS P.835 overall MOS dnsmos p808 DNSMOS P.808 MOS duration Segment duration in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy