Wav2Vec2 XLS R 1B Facebook's Wav2Vec2 XLS R counting 1 billion parameters. XLS R is Facebook AI's large scale multilingual pretrained model for speech (the "XLM R for Speech"). It is pretrained on 436k hours of unlabeled speech, including VoxPopuli, MLS, CommonVoice, BABEL, and VoxLingua107. It uses the wav2vec 2.0 objective, in 128 languages. When using the model make sure that your speech input is sampled at 16kHz. Note : This model should be fine tuned on a downstream task, like Automatic Speech Recognition, Translation, or Classification. Check out this blog for more information about ASR. XLS R Paper Abstract This paper presents XLS R, a large scale model for cross lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on 436K hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work. Our evaluation covers a wide range of tasks, domains, data regimes and languages, both high and low resource. On the CoVoST 2 speech translation benchmark, we improve the previous state of the art by an average of 7.4 BLEU over 21 translation directions into English. For speech…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy