Model card for beitv2 base patch16 224.in1k ft in22k A BEiT v2 image classification model. Trained on ImageNet 1k with self supervised masked image modelling (MIM) using a VQ KD encoder as a visual tokenizer (via OpenAI CLIP B/16 teacher). Fine tuned on ImageNet 22k. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 102.6 GMACs: 17.6 Activations (M): 23.9 Image size: 224 x 224 Papers: BEiT v2: Masked Image Modeling with Vector Quantized Visual Tokenizers: https://arxiv.org/abs/2208.06366 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Dataset: ImageNet 22k Original: https://github.com/microsoft/unilm/tree/master/beit2 Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy