SPEAR Large (speech + general audio) This is the SPEAR Large dual domain (speech + general audio) model. The model adopts a Zipformer backbone with 327M parameters consisting of 112 Zipformer stacks. It generates 1024 dimensional representations at approximately 50~Hz. This model was pre trained on 97k hours of mixture data of English speech and general audio, among which 84k hours are speech data, and the remaining 13k hours are general audio data. It achieves high performance on SUPERB benchmark (for speech representation evaluation) and on HEAR benchmark (for audio representation evaluation). The speech data consists of the following datasets: Dataset Duration (hours) Libriheavy ~50k Gigaspeech ~10k VoxPopuli (en) ~24k The audio data consists of the following datasets: Dataset Duration (hours) AudioSet ~5k Freesound ~2.8k Music4all ~1k VGGSound ~0.5k MTG Jamendo ~3.8k Note : The model is pretrained on 16kHz sampled speech/audio data. When using the model make sure that your input is also sampled at 16kHz. Paper Authors: Xiaoyu Yang, Yifan Yang, Zengrui Jin, Ziyun Cui, Wen Wu, Baoxiang Li, Chao Zhang, Phil Woodland Abstract Self Supervised Learning (SSL) excels at learning generi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy