Model card for vit base patch16 384.augreg in21k ft in1k A Vision Transformer (ViT) image classification model. Trained on ImageNet 21k and fine tuned on ImageNet 1k (with additional augmentation and regularization) in JAX by paper authors, ported to PyTorch by Ross Wightman. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 86.9 GMACs: 49.4 Activations (M): 48.3 Image size: 384 x 384 Papers: How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers: https://arxiv.org/abs/2106.10270 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Dataset: ImageNet 1k Pretrain Dataset: ImageNet 21k Original: https://github.com/google research/vision transformer Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy