Dhivehi Synthetic Voice and Speech Augmentation Dataset This dataset is a multi speaker dataset containing 1.26 million synthetic audio samples (~2,627 hours total). Each sample pairs a Dhivehi sentence with an augmented waveform, created through controlled synthesis, voice cloning, and heavy acoustic perturbations. The dataset was generated to enable ASR, TTS, and voice representation research in low resource Dhivehi, focusing on robustness across pronunciation, prosody, and timbre variance. Process Base text source: sentences from a Dhivehi news corpus. TTS model: speech model fine tuned for Dhivehi phonetics. Voice cloning: reference recordings used to condition synthetic speakers. Augmentations: speed & tempo variation dynamic range compression pitch shifting (± semitones) formant warping & spectral noise reverb & background mix in (random: 50 80 samples added to each subset) pronunciation drift simulation Generation time: ~36 hours of continuous synthesis. Sampling rate: 16 kHz PCM WAV. Each row in the metadata includes: Field Description audio Path to .wav file sentence Dhivehi text string speaker id Original speaker tag (e.g. fh 00) subset id Merged canonical speaker (e.g. f…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy