Multilingual MFA Aligned Speech Dataset A large scale multilingual speech dataset with word level and phoneme level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes. Features Column Type Description audio Audio Audio waveform at 16kHz transcript string Text transcription phoneme sequence string Phoneme sequence with spaces between words words list Word level alignments: [{word, start, end}, ...] phonemes list Phoneme level alignments: [{phoneme, start, end}, ...] source string Original dataset source (e.g., voxpopuli, common voice) Languages & Statistics Language Config Hours Samples Sources English english TBD TBD Common Voice, VoxPopuli, GigaSpeech, Emilia, Genshin Voice, Gemini Speech German german TBD TBD Multilingual LibriSpeech, Emilia French french TBD TBD French Game Voice, Multilingual LibriSpeech, Wolof French ASR Spanish spanish TBD TBD CML TTS, LibriVox, TEDx Spanish Russian russi…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy