GigaAM v3 GigaAM v3 is a Conformer based foundation model with 220–240M parameters, pretrained on diverse Russian speech data using the HuBERT CTC objective. It is the third generation of the GigaAM family and provides state of the art performance on Russian ASR across a wide range of domains. GigaAM v3 includes the following model variants: ssl — self supervised HuBERT–CTC encoder pre trained on 700,000 hours of Russian speech ctc — ASR model fine tuned with a CTC decoder rnnt — ASR model fine tuned with an RNN T decoder e2e ctc — end to end CTC model with punctuation and text normalization e2e rnnt — end to end RNN T model with punctuation and text normalization GigaAM v3 training incorporates new internal datasets: callcenter conversations, speech with background music, natural speech, and speech with atypical characteristics. the models perform on average 30% better on these new domains, while maintaining the same quality as previous GigaAM generations on public benchmarks. The table below reports the Word Error Rate (%) for GigaAM v3 and other existing models over diverse domains. Set Name V3 CTC V3 RNNT T One + LM Whisper : : : : : Open Datasets 3.0 2.6 5.7 12.0 Golos Farfiel…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy