EAT base (Epoch 30, Fine tuned Checkpoint) This is the fine tuned version of the EAT base (Epoch 30, Pre trained Checkpoint), further trained on the AS 2M dataset. Compared to the pre trained model, this version provides enhanced audio representations and typically yields better performance in downstream audio understanding tasks such as classification and captioning. For more details on the EAT framework, please refer to the GitHub repository and our paper EAT: Self Supervised Pre Training with Efficient Audio Transformer. 🔧 Usage You can load and use the model for feature extraction directly via Hugging Face Transformers: 📌 Notes The model supports both frame level (\~50Hz) and utterance level (CLS token) representations. See the feature extraction guide for detailed instructions. 📚 Citation If you find this model useful, please consider citing our paper: bibtex @article{chen2024eat, title={EAT: Self supervised pre training with efficient audio transformer}, author={Chen, Wenxi and Liang, Yuzhe and Ma, Ziyang and Zheng, Zhisheng and Chen, Xie}, journal={arXiv preprint arXiv:2401.03497}, year={2024} }
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy