This model is a fine tuned version of facebook/wav2vec2 xls r 300m on the OPENSLR SLR53 bengali dataset. It achieves the following results on the evaluation set. Without language model : WER: 0.21726385291857586 CER: 0.04725010353701041 With 5 gram language model trained on 30M sentences randomly chosen from AI4Bharat IndicCorp dataset : WER: 0.15322879016421437 CER: 0.03413696666806267 Note : 5% of a total 10935 samples have been used for evaluation. Evaluation set has 10935 examples which was not part of training training was done on first 95% and eval was done on last 5%. Training was stopped after 180k steps. Output predictions are available under files section. Training hyperparameters The following hyperparameters were used during training: dataset name="openslr" model name or path="facebook/wav2vec2 xls r 300m" dataset config name="SLR53" output dir="./wav2vec2 xls r 300m bengali" overwrite output dir num train epochs="50" per device train batch size="32" per device eval batch size="32" gradient accumulation steps="1" learning rate="7.5e 5" warmup steps="2000" length column name="input length" evaluation strategy="steps" text column name="sentence" chars to ignore , ? . ! \…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy