Web SSL DINO ViT 300M: 2B MetaCLIP data, 224 Resolution A 300 million parameter Vision Transformer (ViT) trained with DINOv2 self supervised learning on web scale image data without language supervision. Introduced in "Scaling Language Free Visual Representation Learning" (Fan et al., 2025). Model Details Architecture : ViT (1536 width, 40 depth, 24 heads) Parameters : 300M Resolution : 224×224 pixels Training : Self supervised Web DINO on 2B image samples from MetaCLIP web data Model Descriptions Web SSL DINO 300M is a 300 million parameter Vision Transformer model trained using self supervised learning on 2 billion web images without language supervision. This model demonstrates that pure visual learning, when scaled appropriately, can match or exceed the performance of language supervised models like CLIP across various vision tasks. Usage Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy