JaiTTS F5TTS: Thai Voice Cloning Model Research Prototype JaiTTS F5TTS is a non autoregressive JaiTTS voice cloning model based on F5 TTS. It targets Thai zero shot voice cloning. This model was presented in the paper JaiTTS: A Thai Voice Cloning Model. Research prototype: JaiTTS F5TTS is one of our experimental variants within the JaiTTS project. It is released for research and benchmarking only. Highlights F5 TTS based non autoregressive voice cloning for Thai Duration predictor for improved pacing and intelligibility Fast synthesis with Real Time Factor (RTF) below 0.2 Duration Modeling The original F5 TTS duration estimate uses a UTF 8 byte ratio formula. This is brittle for Thai and mixed script input because Thai characters, English words, Arabic numerals, and punctuation do not have a consistent byte to pronunciation relationship. In practice, the mismatch can produce rushed, compressed, or unstable speech. We address this with an XLM R based neural duration predictor that estimates target duration from text more robustly than the UTF 8 byte ratio baseline. The data used to train and evaluate the duration predictor is sampled from the JaiTTS v1.0 training set. Duration Predi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy