TTS Pretrain Clones (3M) 2,967,779 clone utterances across 2971 English speakers. Sample rate: 44.1 kHz, WAV in Parquet Generated by echo tts synthesizing English text on speaker latents derived from Qwen3 TTS VoiceDesign base speakers. Per speaker: 10 voice clone latents × 100 texts. The first utterance of each speaker (row 0) is published separately in the companion refs set. Coverage: speakers 1 60 + 61 (partial, 749 rows) + 91 3000. Thirty speakers (61's tail + 62 90) are… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab/tts pretrain clones 3m.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy