Whisper Tamil Small This model is a fine tuned version of openai/whisper small on the Tamil data available from multiple publicly available ASR corpuses. It has been fine tuned as a part of the Whisper fine tuning sprint. NOTE: The code used to train this model is available for re use in the whisper finetune repository. Usage In order to evaluate this model on an entire dataset, the evaluation codes available in the whisper finetune repository can be used. The same repository also provides the scripts for faster inference using whisper jax. In order to infer a single audio file using this model, the following code snippet can be used: For faster inference of whisper models, the whisper jax library can be used. Please follow the necessary installation steps as mentioned here, before using the following code snippet: Training and evaluation data Training Data: IISc MILE Tamil ASR Corpus ULCA ASR Corpus Shrutilipi ASR Corpus Microsoft Speech Corpus (Indian Languages) Google/Fleurs Train+Dev set Babel ASR Corpus Evaluation Data: Microsoft Speech Corpus (Indian Languages) Test Set Google/Fleurs Test Set IISc MILE Test Set Babel Test Set Training hyperparameters The following hyperparame…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy