Malaysian Emilia An Extensive, Multilingual, and Diverse Speech Dataset for Large Scale Malaysian and Singaporean Speech Generation, where originally from Emilia. We are improving Malaysian Emilia due to https://github.com/open mmlab/Amphion/issues/436, check out mesolitica/Malaysian Emilia v2 Dataset Clone and Extract We upload as split zip files so you can clone and extract distributedly, Malaysian Cartoons 1. Originally from malaysia ai/malaysian cartoons youtube, total 20.8k hours. 3. 774.5 hours after processed, 332187 audio files, malaysian cartoon.zip Malaysian Youtube 1. Originally from malaysia ai/malaysian youtube, total 18.7k hours. 2. 3168.8 hours after processed, 1014187 audio files, filtered processed 0 0.zip 3. post cleaned to 24k and 44k samples rate at mesolitica/Malaysian Emilia annotated Malaysian Podcast 1. Originally from malaysia ai/malaysian podcast youtube, total 2.2k hours. 2. 622.8 hours after processed, 213164 audio files, malaysian podcast processed.zip 3. post cleaned to 24k and 44k sample rates at mesolitica/Malaysian Emilia annotated Singaporean Podcast 1. Originally from malaysia ai/singaporean podcast youtube, total 1.2k hours. 2. 175.9 hours after…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy