Model card for vit base patch16 1024 128.audiomae as2m ft as20k A Vision Transformer (ViT) for audio. Pretrained on AudioSet 2M with Self Supervised Masked Autoencoder (MAE) method, and fine tuned on AudioSet 20k. This is a port of AudioMAE ViT B/16 weights for usage with timm . The naming convention is adopted from other timm 's ViT models. See the original repo here: https://github.com/facebookresearch/AudioMAE For the AudioSet 2M pre trained checkpoint (without Audioset 20k fine tuning), see https://huggingface.co/gaunernst/vit base patch16 1024 128.audiomae as2m Model Details Model Type: Audio classification / feature backbone Papers: Masked Autoencoders that Listen: https://arxiv.org/abs/2207.06405 Pretrain Dataset: AudioSet 2M Original: https://github.com/facebookresearch/AudioMAE Model Usage Audio Classification and Embeddings Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy