Voxtral 4B TTS 2603 Voxtral TTS is a frontier, open weights text to speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY NC 4, which is the license that the model inherits. For more details, see our: 🔊 Demo ✍️ Blog post 🔬 Research Paper Key Features Voxtral TTS delivers enterprise grade text to speech for production voice agents, with the following capabilities: Realistic, expressive speech with natural prosody and emotional range across 9 major languages, with support for diverse dialects Text to Speech generation with 20 preset voices and easy adaptation to new voices Multilingual support : English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, and Hindi Very low latency with fast time to first audio, plus streaming and batch inference support 24 kHz audio output in WAV, PCM, FLAC, MP3, AAC, and Opus formats Production ready performance for high throughput, real time voice agent workflows [!Tip] For voice customization, visit our AI Studio. Use Cases Customer support and call center infrastructure. Financial service…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy