Whisper Whisper is a state of the art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on 5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero shot setting. Whisper large v3 turbo is a finetuned version of a pruned Whisper large v3. In other words, it's the exact same model, except that the number of decoding layers have reduced from 32 to 4. As a result, the model is way faster, at the expense of a minor quality degradation. You can find more details about it in this GitHub discussion. Disclaimer : Content for this model card has partly been written by the 🤗 Hugging Face team, and partly copied and pasted from the original model card. Usage Whisper large v3 turbo is supported in Hugging Face 🤗 Transformers. To run the model, first install the Transformers library. For this example, we'll also install 🤗 Datasets to load toy audio dataset from the Hugging Face Hub, and 🤗 Accelerate to reduce the model loading time: The model can be used with the pipeline class to transcribe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy