YODAS2 Sidon Overview This dataset is a cleansed version of YODAS 2 with Sidon speech restoration mode for Speech Synthesis and Spoken Language Modeling . YODAS 2 is a massive, multilingual YouTube derived dataset. We have applied the Sidon restoration model to remove background noise and enhance audio quality, making it suitable for high quality generation tasks. We resampled original sidon output to 24kHz due to a storage constraints. The dataset is provided in WebDataset format for efficient large scale training. Source : YODAS 2 (YouTube Oriented Dataset for Audio Visual Speech) Format : WebDataset ( .tar.gz shards) License : CC BY 3.0 Dataset Structure Each sample in the dataset contains: flac — audio file (24 kHz, single channel, restored) metadata.json (optional) — metadata including language, YouTube video ID, and transcription Example (inside a .tar shard): How to Use With 🤗 Datasets You can load the WebDataset directly with Hugging Face’s datasets library: Replace subset with the desired subset. Citation If you use this dataset, please cite Sidon and the original YODAS paper: License This dataset is released under CC BY 3.0. Acknowledgements Original data : YODAS2
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy