π News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025! Open Whisper style Speech Model (OWSM) is the first fully open Whisper style speech foundation model. It reproduces and advances OpenAI's Whisper style training using publicly available data and open source toolkits. The code, pre trained model weights, and training logs are publicly released to promote open science in speech foundation models. OWSM CTC (Peng et al., ACL 2024) is a novel encoder only speech foundation model based on hierarchical multi task self conditioned CTC. It supports multilingual speech recognition, speech translation, and language identification within a single non autoregressive model. OWSM CTC v4 is trained for three epochs on 320k hours of public audio data covering multilingual speech recognition, any to any speech translation, and language identification. The newly curated data are publicly released: https://huggingface.co/datasets/espnet/yodas owsmv4 To use the pre trained model, please install espnet and espnet model zoo . The requirements are: The recipe can be found in ESPnet: https://github.com/espnet/espnet/tree/master/egs2/owsm ctc v4/s2t1 Example script for batchedβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy