Wav2Vec2 Large XLSR 53 Swedish Fine tuned facebook/wav2vec2 large xlsr 53 in Swedish using the NST Swedish Dictation. When using this model, make sure that your speech input is sampled at 16kHz. Note: We recommend using our newer model wav2vec2 large voxrex swedish for the best performance. Usage The model can be used directly (without a language model) as follows: Evaluation The model can be evaluated as follows on the Swedish test data of Common Voice. WER : 14.298610% CER : 4.925294% Training First the XLSR model was further pre trained for 50 epochs with a corpus consisting of 1000 hours spoken Swedish from various radio stations. Secondly NST Swedish Dictation was used for fine tuning as well as Common Voice. Lastly only Common Voice dataset was used for final finetuning. The Fairseq scripts were used.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy