Model card for eva02 small patch14 336.mim in22k ft in1k An EVA02 image classification model. Pretrained on ImageNet 22k with masked image modeling (using EVA CLIP as a MIM teacher) and fine tuned on ImageNet 1k by paper authors. EVA 02 models are vision transformers with mean pooling, SwiGLU, Rotary Position Embeddings (ROPE), and extra LN in MLP (for Base & Large). NOTE: timm checkpoints are float32 for consistency with other models. Original checkpoints are float16 or bfloat16 in some cases, see originals if that's preferred. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 22.1 GMACs: 15.5 Activations (M): 54.3 Image size: 336 x 336 Papers: EVA 02: A Visual Representation for Neon Genesis: https://arxiv.org/abs/2303.11331 EVA CLIP: Improved Training Techniques for CLIP at Scale: https://arxiv.org/abs/2303.15389 Original: https://github.com/baaivision/EVA https://huggingface.co/Yuxin CV/EVA 02 Pretrain Dataset: ImageNet 22k Dataset: ImageNet 1k Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. model top1 top5 param count img size eva02 large pat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy