Model card for vit base patch16 224.orig in21k A Vision Transformer (ViT) image classification model. Pretrained on ImageNet 21k in JAX by paper authors, ported to PyTorch by Ross Wightman. This model does not have a classification head, useful for features and fine tune only. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 85.8 GMACs: 16.9 Activations (M): 16.5 Image size: 224 x 224 Papers: An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Dataset: ImageNet 21k Original: https://github.com/google research/vision transformer Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy