SSLAM AudioSet 2M Finetuned (ViT Base, mAP:50.2) This repository provides an SSLAM checkpoint in Hugging Face Transformers format. You can use it to extract audio embeddings or to obtain sound event class labels for event detection. The implementation follows the EAT code path with SSLAM AudioSet 2M finetuned weights. ๐ง Usage You can load and use the model for feature extraction directly via Hugging Face Transformers: ๐ Notes See the feature extraction guide for more instructions. ๐ Acknowledgments This repository builds on the EAT implementation for Hugging Face models. We remap SSLAM weights to that interface. We are not affiliated with the EAT authors. All credit for the original implementation belongs to them. ๐ Citation If you find our work useful, please cite it as: Please also cite EAT:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy