Model card for vit base patch16 clip 224.laion2b ft in12k in1k A Vision Transformer (ViT) image classification model. Pretrained on LAION 2B image text pairs using OpenCLIP. Fine tuned on ImageNet 12k and then ImageNet 1k in timm . See recipes in Reproducible scaling laws. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 86.6 GMACs: 16.9 Activations (M): 16.5 Image size: 224 x 224 Papers: OpenCLIP: https://github.com/mlfoundations/open clip Reproducible scaling laws for contrastive language image learning: https://arxiv.org/abs/2212.07143 LAION 5B: An open large scale dataset for training next generation image text models: https://arxiv.org/abs/2210.08402 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Dataset: ImageNet 1k Pretrain Dataset: LAION 2B ImageNet 12k Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy