Pretrained Model Fine tuned on Multilingual Pretrained Model CLSRIL 23. The original fairseq checkpoint is present here. When using this model, make sure that your speech input is sampled at 16kHz. Note: The result from this model is without a language model so you may witness a higher WER in some cases. Dataset This model was trained on 4200 hours of Hindi Labelled Data. The labelled data is not present in public domain as of now. Training Script Models were trained using experimental platform setup by Vakyansh team at Ekstep. Here is the training repository. In case you want to explore training logs on wandb they are here. Colab Demo Usage The model can be used directly (without a language model) as follows: Evaluation The model can be evaluated as follows on the hindi test data of Common Voice. Test Result : 53.64 % Colab Evaluation Credits Thanks to Ekstep Foundation for making this possible. The vakyansh team will be open sourcing speech models in all the Indic Languages.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy