Model card for eva02 large patch14 448.mim m38m ft in22k An EVA02 image classification model. Pretrained on Merged 38M (IN 22K, CC12M, CC3M, COCO (train), ADE20K (train), Object365, and OpenImages) with masked image modeling (using EVA CLIP as a MIM teacher) and fine tuned on ImageNet 22k by paper authors. EVA 02 models are vision transformers with mean pooling, SwiGLU, Rotary Position Embeddings (ROPE), and extra LN in MLP (for Base & Large). NOTE: timm checkpoints are float32 for consistency with other models. Original checkpoints are float16 or bfloat16 in some cases, see originals if that's preferred. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 326.4 GMACs: 362.4 Activations (M): 690.0 Image size: 448 x 448 Papers: EVA 02: A Visual Representation for Neon Genesis: https://arxiv.org/abs/2303.11331 EVA CLIP: Improved Training Techniques for CLIP at Scale: https://arxiv.org/abs/2303.15389 Original: https://github.com/baaivision/EVA https://huggingface.co/Yuxin CV/EVA 02 Pretrain Dataset: ImageNet 22k Dataset: ImageNet 22k Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy