SPEAR XLarge (speech + general audio) Recommended: This is the original SPEAR XLarge v1 checkpoint. For new projects, we recommend using the ICML 2026 accepted XLarge v2 model : marcoyang/spear xlarge speech audio v2. XLarge v2 is aligned with the accepted ICML 2026 version and includes additional robustness improvements for complex acoustic scenes through token mixing. This is the SPEAR XLarge v1 dual domain (speech + general audio) model. The model adopts a Zipformer backbone with 597M parameters consisting of 13 Zipformer stacks. It generates 1280 dimensional representations at approximately 50~Hz. The model was pre trained on 197k hours of mixture data of English speech and general audio, among which 184k hours are speech data, and the rest 13k hours are general audio data. It achieves state of the art performance on the SUPERB benchmark and competitive performance on HEAR. For the latest SPEAR XLarge model with stronger complex scene robustness, please use XLarge v2. The speech data consists of the following datasets: Dataset Duration (hours) Libriheavy ~50k Gigaspeech ~10k VoxPopuli (en) ~24k yodas granary ~100k The audio data consists of the following datasets: Dataset Durat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy