SEW D tiny SEW D by ASAPP Research The base model pretrained on 16kHz sampled speech audio. When using the model make sure that your speech input is also sampled at 16Khz. Note that this model should be fine tuned on a downstream task, like Automatic Speech Recognition, Speaker Identification, Intent Classification, Emotion Recognition, etc... Paper: Performance Efficiency Trade offs in Unsupervised Pre training for Speech Recognition Authors: Felix Wu, Kwangyoun Kim, Jing Pan, Kyu Han, Kilian Q. Weinberger, Yoav Artzi Abstract This paper is a study of performance efficiency trade offs in pre trained models for automatic speech recognition (ASR). We focus on wav2vec 2.0, and formalize several architecture designs that influence both the model performance and its efficiency. Putting together all our observations, we introduce SEW (Squeezed and Efficient Wav2vec), a pre trained model architecture with significant improvements along both performance and efficiency dimensions across a variety of training setups. For example, under the 100h 960h semi supervised setup on LibriSpeech, SEW achieves a 1.9x inference speedup compared to wav2vec 2.0, with a 13.5% relative reduction in word er…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy