🦜 VieNeu TTS v3 Turbo Overview VieNeu TTS v3 Turbo is the next generation of Vietnamese TTS — 48 kHz high fidelity speech, instant voice cloning , built in multi speaker default voices , inline emotion cues , and seamless bilingual (En–Vi) code switching . It is a pure PyTorch engine running on both GPU and CPU , using the MOSS Audio Tokenizer Nano codec. [!NOTE] Early access. v3 Turbo is released for preview . It is fast and natural, but some features (notably the emotion cues) are still experimental . The full v3 release is coming in the next few weeks. [!IMPORTANT] What's new in v3: 48 kHz audio — a big jump in fidelity over v2 (24 kHz). Built in default voices — each default speaker is addressed by a dedicated speaker token + fixed reference, so the voice is stable and consistent with no reference clip needed. Emotion / non verbal cues (experimental) — drop [cười] , [thở dài] , [hắng giọng] straight into your text. Batched generation — synthesize many chunks at once (batch size up to 32), including a multi speaker conversation mode that batches the whole script regardless of speaker. Instant Voice Cloning — still clones a voice from just 3–5 seconds of audio (cloning is availa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy