Wav2Vec2 XLS R 300M Facebook's Wav2Vec2 XLS R counting 300 million parameters. XLS R is Facebook AI's large scale multilingual pretrained model for speech (the "XLM R for Speech"). It is pretrained on 436k hours of unlabeled speech, including VoxPopuli, MLS, CommonVoice, BABEL, and VoxLingua107. It uses the wav2vec 2.0 objective, in 128 languages. When using the model make sure that your speech input is sampled at 16kHz. Note : This model should be fine tuned on a downstream task, like Automatic Speech Recognition, Translation, or Classification. Check out this blog for more information about ASR. XLS R Paper Authors: Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, Alexei Baevski, Alexis Conneau, Michael Auli Abstract This paper presents XLS R, a large scale model for cross lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on 436K hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work. Our evaluation covers a wide range of tasks, domains, data regimes and la…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy