Qwen3 TTS   🤗 Hugging Face      🤖 ModelScope      📑 Blog      📑 Paper      💻 GitHub We release Qwen3 TTS , a series of powerful speech generation models developed by Qwen, offering comprehensive support for voice cloning, voice design, ultra high quality human like speech generation, and natural language based voice control. Overview Qwen3 TTS covers 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) as well as multiple dialectal voice profiles. Key features: Powerful Speech Representation : Powered by the self developed Qwen3 TTS Tokenizer 12Hz, it achieves efficient acoustic compression and high dimensional semantic modeling. Universal End to End Architecture : Utilizing a discrete multi codebook LM architecture to bypass traditional information bottlenecks. Extreme Low Latency Streaming Generation : Supports streaming generation with end to end synthesis latency as low as 97ms. Intelligent Voice Control : Supports speech generation driven by natural language instructions for flexible control over timbre, emotion, and prosody. Quickstart Env…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy