FireRedASR2S A SOTA Industrial Grade All in One ASR System [[Code]](https://github.com/FireRedTeam/FireRedASR2S) [[Paper]](https://arxiv.org/pdf/2501.14350) [[Model]](https://huggingface.co/FireRedTeam) [[Blog]](https://fireredteam.github.io/demos/firered asr/) [[Demo]](https://huggingface.co/spaces/FireRedTeam/FireRedASR) FireRedASR2S is a state of the art (SOTA), industrial grade, all in one ASR system with ASR, VAD, LID, and Punc modules. All modules achieve SOTA performance: FireRedASR2 : Automatic Speech Recognition (ASR) supporting Chinese (Mandarin, 20+ dialects/accents), English, code switching, and singing lyrics recognition. 2.89% average CER on Mandarin (4 test sets), 11.55% on Chinese dialects (19 test sets), outperforming Doubao ASR, Qwen3 ASR 1.7B, Fun ASR, and Fun ASR Nano 2512. FireRedASR2 AED also supports word level timestamps and confidence scores. FireRedVAD : Voice Activity Detection (VAD) supporting speech/singing/music in 100+ languages. 97.57% F1, outperforming Silero VAD, TEN VAD, and FunASR VAD. Supports non streaming/streaming VAD and Audio Event Detection. FireRedLID : Spoken Language Identification (LID) supporting 100+ languages and 20+ Chinese dialect…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy