Model description Our models use wav2vec2 architecture, pre trained on 13k hours of Vietnamese youtube audio (un label data) and fine tuned on 250 hours labeled of VLSP ASR dataset on 16kHz sampled speech audio. You can find more description here Benchmark WER result on VLSP T1 testset: base model large model without LM 8.66 6.90 with 5 grams LM 6.53 5.32 Usage Model Parameters License The ASR model parameters are made available for non commercial use only, under the terms of the Creative Commons Attribution NonCommercial 4.0 International (CC BY NC 4.0) license. You can find details at: https://creativecommons.org/licenses/by nc/4.0/legalcode Contact nguyenvulebinh@gmail.com
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy