Qwen3 Omni Overview Introduction Qwen3 Omni is the natively end to end multilingual omni modal foundation models. It processes text, images, audio, and video, and delivers real time streaming responses in both text and natural speech. We introduce several architectural upgrades to improve performance and efficiency. Key features: State of the art across modalities : Early text first pretraining and mixed multimodal training provide native multimodal support. While achieving strong audio and audio video results, unimodal text and image performance does not regress. Reaches SOTA on 22 of 36 audio/video benchmarks and open source SOTA on 32 of 36; ASR, audio understanding, and voice conversation performance is comparable to Gemini 2.5 Pro. Multilingual : Supports 119 text languages, 19 speech input languages, and 10 speech output languages. Speech Input : English, Chinese, Korean, Japanese, German, Russian, Italian, French, Spanish, Portuguese, Malay, Dutch, Indonesian, Turkish, Vietnamese, Cantonese, Arabic, Urdu. Speech Output : English, Chinese, French, German, Russian, Italian, Spanish, Portuguese, Japanese, Korean. Novel Architecture : MoE based Thinker–Talker design with AuT pre…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy