Dataset Card for MultiLingual LibriSpeech Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages How to use Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: MultiLingual LibriSpeech ASR corpus Repository: [Needs More Information] Paper: MLS: A Large Scale Multilingual Dataset for Speech Research Leaderboard: 🤗 Autoevaluate Leaderboard Dataset Summary This is a streamable version of the Multilingual LibriSpeech (MLS) dataset. The data archives were restructured from the original ones from OpenSLR to make it easier to stream. MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages English, German, Dutch, Spanish, French, Italian, Portuguese, Polish. It includes about 44.5K hours of English and a total of about 6K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy