MLS Sidon Overview This dataset is a cleansed version of Multilingual LibriSpeech (MLS) with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling . The dataset is provided in WebDataset format for efficient large scale training. Source : Multilingual LibriSpeech Languages : English, German, French, Spanish, Italian, Polish, Dutch, Portuguese Format : WebDataset ( .tar shards) License : CC BY 4.0 Dataset Structure Each sample in the dataset contains: flac — audio file (48 kHz, single channel) metadata.json (optional) — metadata including language, speaker ID, and original MLS reference Example (inside a .tar shard): How to Use With 🤗 Datasets You can load the WebDataset directly with Hugging Face’s datasets library: Replace language with the language (e.g., english , german ). Citation If you use this dataset, please cite Sidon and the original MLS paper: License This dataset is released under CC BY 4.0. Acknowledgements Original data : Multilingual LibriSpeech (MLS)
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy