SpeechT5 (ASR task) SpeechT5 model fine tuned for automatic speech recognition (speech to text) on LibriSpeech. This model was introduced in SpeechT5: Unified Modal Encoder Decoder Pre Training for Spoken Language Processing by Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei. SpeechT5 was first released in this repository, original weights. The license used is MIT. Disclaimer: The team releasing SpeechT5 did not write a model card for this model so this model card has been written by the Hugging Face team. Model Description Motivated by the success of T5 (Text To Text Transfer Transformer) in pre trained natural language processing models, we propose a unified modal SpeechT5 framework that explores the encoder decoder pre training for self supervised speech/text representation learning. The SpeechT5 framework consists of a shared encoder decoder network and six modal specific (speech/text) pre/post nets. After preprocessing the input speech/text through the pre nets, the shared encoder decoder network models the sequence to sequence transformation, and then the post nets generate the outpu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy