LargeScaleASR: 25,000 hours of transcribed and heterogeneous English speech recognition data for research and commercial use. The full details are available in the paper. Made of 6 subsets: 1. large contains 25,000 hours of read / spontaneous and clean / noisy transcribed speech. 2. medium contains 2,500 hours of read / spontaneous and clean / noisy transcribed speech. 3. small contains 250 hours of read / spontaneous and clean / noisy transcribed speech. 4. clean contains 13,000 hours of read / spontaneous transcribed speech. YODA and People's Speech data are excluded from this subset as, despite data curation, some errors remain in the transcriptions. 5. dev contains 15 hours (more details in the next section). 6. test contains 21 hours (more details in the next section). The large split requires 4TB of storage (including HuggingFace extraction). The shards only are 2TB. Example: Training recipe A full conformer ASR training recipe is available here. Data description (Following information are directly copy pasted from the SpeechBrain data preparation README) TLS is a mix of 5 existing dataset with permissive licences. The way it is mixed is described in the following table: Data…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy