Qwen3 ASR (Transformers native) Overview The Qwen3 ASR family includes Qwen3 ASR 1.7B and Qwen3 ASR 0.6B , which support language identification and ASR for 52 languages and dialects. Both leverage large scale speech training data and the strong audio understanding capability of their foundation model, Qwen3 Omni. The 1.7B version achieves state of the art performance among open source ASR models and is competitive with the strongest proprietary commercial APIs. Key features: All in one: Supports language identification and speech recognition for 30 languages and 22 Chinese dialects, including English accents from multiple countries and regions. Excellent and Fast: High quality and robust recognition under complex acoustic environments. Qwen3 ASR 0.6B reaches 2000× throughput at a concurrency of 128. Both models support streaming/offline unified inference with a single model and handle long audio. Forced Alignment: Qwen3 ForcedAligner 0.6B supports timestamp prediction for arbitrary units within up to 5 minutes of speech in 11 languages, surpassing E2E based forced alignment models in accuracy. Model Architecture Available Checkpoints Model Supported Languages Supported Dialects In…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy