Swahili Automatic Speech Recognition (ASR) Model details The Swahili ASR is an end to end automatic speech recognition system that was finetuned on the Common Voice Corpus 11.0 Swahili dataset. This repository provides the necessary tools to perform ASR using this model, allowing for high quality speech to text conversions in Swahili. Example Usage Here's an example of how you can use this model for speech to text conversion: EVAL LOSS EVAL WER EVAL RUNTIME EVAL SAMPLES PER SECOND EVAL STEPS PER SECOND EPOCH 0.345414400100708 0.2602372795622284 578.4006 17.701 2.213 4.17 Intended Use This model is intended for any application requiring Swahili speech to text conversion, including but not limited to transcription services, voice assistants, and accessibility technology. It can be particularly beneficial in any context where demographic metadata (age, sex, accent) is significant, as these features have been taken into account during training. Dataset The model was trained on the Common Voice Corpus 11.0 Swahili dataset, which consists of unique MP3 files and corresponding text files, totaling 16,413 validated hours. Additionally, much of the dataset includes valuable demographic meta…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy