Wav2vec 2.0 large VoxRex Swedish (C) Finetuned version of KBs VoxRex large model using Swedish radio broadcasts, NST and Common Voice data. Evalutation without a language model gives the following: WER for NST + Common Voice test set (2% of total sentences) is 2.5% . WER for Common Voice test set is 8.49% directly and 7.37% with a 4 gram language model. When using this model, make sure that your speech input is sampled at 16kHz. Update 2022 01 10: Updated to VoxRex C version. Update 2022 05 16: Paper is is here. Performance\ Chart shows performance without the additional 20k steps of Common Voice fine tuning Training This model has been fine tuned for 120000 updates on NST + CommonVoice and then for an additional 20000 updates on CommonVoice only. The additional fine tuning on CommonVoice hurts performance on the NST+CommonVoice test set somewhat and, unsurprisingly, improves it on the CommonVoice test set. It seems to perform generally better though [citation needed] . Usage The model can be used directly (without a language model) as follows: Citation https://arxiv.org/abs/2205.03026
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy