Model card for vit small patch14 dinov2.lvd142m A Vision Transformer (ViT) image feature model. Pretrained on LVD 142M with self supervised DINOv2 method. Model Details Model Type: Image classification / feature backbone Model Stats: Params (M): 22.1 GMACs: 46.8 Activations (M): 198.8 Image size: 518 x 518 Papers: DINOv2: Learning Robust Visual Features without Supervision: https://arxiv.org/abs/2304.07193 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale: https://arxiv.org/abs/2010.11929v2 Original: https://github.com/facebookresearch/dinov2 Pretrain Dataset: LVD 142M Model Usage Image Classification Image Embeddings Model Comparison Explore the dataset and runtime metrics of this model in timm model results. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy