A CLIP (Contrastive Language Image Pre training) model trained on DFN 5B. Data Filtering Networks (DFNs) are small networks used to automatically filter large pools of uncurated data. This model was trained on 5B images that were filtered from a pool of 43B uncurated image text pairs (12.8B image text pairs from CommonPool 12.8B + 30B additional public image text pairs). This model has been converted to PyTorch from the original JAX checkpoints from Axlearn (https://github.com/apple/axlearn). These weights are directly usable in OpenCLIP (image + text). Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Dataset: DFN 5b Papers: Data Filtering Networks: https://arxiv.org/abs/2309.17425 Samples Seen: 39B (224 x 224) + 5B (384 x 384) Model Metrics dataset metric : : ImageNet 1k 0.84218 Caltech 101 0.954479 CIFAR 10 0.9879 CIFAR 100 0.9041 CLEVR Counts 0.362467 CLEVR Distance 0.206067 Country211 0.37673 Describable Textures 0.71383 EuroSAT 0.608333 FGVC Aircraft 0.719938 Food 101 0.963129 GTSRB 0.679018 ImageNet Sketch 0.73338 ImageNet v2 0.7837 ImageNet A 0.7992 ImageNet O 0.3785 ImageNet R 0.937633 KITTI Vehicle Distance 0.38256 MNIST 0.8372 ObjectNet 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy