Dataset Card for CML TTS Table of Contents Dataset Description Dataset Summary Supported Tasks Languages How to use Dataset Structure Data Instances Data Fields Data Splits Data Statistics Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: MultiLingual LibriSpeech ASR corpus Repository: CML TTS Dataset Paper: CML TTS A Multilingual Dataset for Speech Synthesis in Low Resource Languages Dataset Summary CML TTS is a recursive acronym for CML Multi Lingual TTS, a Text to Speech (TTS) dataset developed at the Center of Excellence in Artificial Intelligence (CEIA) of the Federal University of Goias (UFG). CML TTS is a dataset comprising audiobooks sourced from the public domain books of Project Gutenberg, read by volunteers from the LibriVox project. The dataset includes recordings in Dutch, German, French, Italian, Polish, Portuguese, and Spanish, all at a sampling rate of 24kHz. The data archives were restructured…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy