Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia Pipe speech data preprocessing pipeline. News 🔥 2025/02/26 : The Emilia Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia Large combines the original 101k hour Emilia dataset (licensed under CC BY NC 4.0 ) with the brand new 114k hour Emilia YODAS dataset (licensed under CC BY 4.0 )!!! 2025/01/27 : We release the extended version of Emilia's paper on arXiv! More experiments and more insights! 2024/12/04 : We present Emilia at the IEEE SLT 2024! 2024/08/28 : Welcome to join Amphion's Discord channel to stay connected and engage with our community! 2024/08/27 : The Emilia dataset is now publicly available! Discover the most extensive and diverse speech generation dataset with 101k hours of in the wild speech data now at HuggingFace or OpenDataLab! 👑👑👑 2024/07/08 : Our preprint paper is now available! 🔥🔥🔥 2024/07/03 : We welcome everyone to check our homepage for our brief introduction for Emilia dataset and our demos! 2024/07/01 : We release of Emilia and Emili…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy