Wav2Vec2 Large Robust finetuned on Librispeech Facebook's Wav2Vec2. This model is a fine tuned version of the wav2vec2 large robust model. It has been pretrained on: Libri Light: open source audio books from the LibriVox project; clean, read out audio data CommonVoice: crowd source collected audio data; read out text snippets Switchboard: telephone speech corpus; noisy telephone data Fisher: conversational telephone speech; noisy telephone data and subsequently been finetuned on 960 hours of Librispeech: open source read out audio data. When using the model make sure that your speech input is also sampled at 16Khz. Paper Robust Wav2Vec2 Authors: Wei Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, Michael Auli Abstract Self supervised learning of speech representations has been a very active research area but most work is focused on a single domain such as read audio books for which there exist large quantities of labeled and unlabeled data. In this paper, we explore more general setups where the domain of the unlabeled data for pre training data differs from the domain of the labeled…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy