EAT base (Epoch 30, Pre trained Checkpoint) This is the pre trained EAT base model at epoch 30, trained on the AS 2M dataset using the EAT framework for audio self supervised learning. It offers efficient feature extraction and can also serve as a strong initialization for fine tuning on a wide range of downstream audio understanding tasks such as classification and captioning. For more details on the EAT framework, please refer to the GitHub repository and our paper EAT: Self Supervised Pre Training with Efficient Audio Transformer. 🔧 Usage You can load and use the model for feature extraction directly via Hugging Face Transformers: 📌 Notes The model supports both frame level (\~50Hz) and utterance level (CLS token) representations. See the feature extraction guide for more instructions. 📚 Citation If you find this model useful, please consider citing our paper: bibtex @article{chen2024eat, title={EAT: Self supervised pre training with efficient audio transformer}, author={Chen, Wenxi and Liang, Yuzhe and Ma, Ziyang and Zheng, Zhisheng and Chen, Xie}, journal={arXiv preprint arXiv:2401.03497}, year={2024} }
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy