Whisper IPA Whisper is a pre trained model for automatic speech recognition (ASR) and speech translation. Fine tuned on 15000 wavs of labelled synthetic IPA data (generated using the goruut 0.6.2 phonemizer), Whisper models demonstrate a strong ability to generalise to many languages, datasets and domains without the need for fine tuning. Whisper was proposed in the paper Robust Speech Recognition via Large Scale Weak Supervision by Alec Radford et al from OpenAI. The original code repository can be found here. Disclaimer : Content for this model card has partly been written by the Hugging Face team, and parts of it were copied and pasted from the original model card. Fine tuning details Fine tuning took 20:44:16 It was trained on 15000 wavs GPU in use was NVIDIA 3090ti with 24GB VRAM Fine tuned on 15000 random wavs from common voice 21 across 70+ languages Model details Whisper is a Transformer based encoder decoder model, also referred to as a sequence to sequence model. It was trained on 680k hours of labelled speech data annotated using large scale weak supervision. The models were trained on either English only data or multilingual data. The English only models were trained on…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy