Model card for beit base patch16 224.in22k ft in22k in1k A BEiT image classification model. Trained on ImageNet 22k with self supervised masked image modelling (MIM) using a DALL E dVAE as visual tokenizer. Fine tuned on ImageNet 22k and then ImageNet 1k. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 86.5 GMACs: 17.6 Activations (M): 23.9 Image size: 224 x 224 Papers: BEiT: BERT Pre Training of Image Transformers: https://arxiv.org/abs/2106.08254 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Dataset: ImageNet 1k Pretrain Dataset: ImageNet 22k Original: https://github.com/microsoft/unilm/tree/master/beit Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy