A CLIP (Contrastive Language Image Pre training) model trained on DFN 5B. Data Filtering Networks (DFNs) are small networks used to automatically filter large pools of uncurated data. This model was trained on 5B images that were filtered from a pool of 43B uncurated image text pairs (12.8B image text pairs from CommonPool 12.8B + 30B additional public image text pairs). This model has been converted to PyTorch from the original JAX checkpoints from Axlearn (https://github.com/apple/axlearn). These weights are directly usable in OpenCLIP (image + text). Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Dataset: DFN 5b Papers: Data Filtering Networks: https://arxiv.org/abs/2309.17425 Samples Seen: 39B Model Metrics Eval Dataset Metric : : ImageNet 1k 0.8344 Caltech 101 0.954935 CIFAR 10 0.9878 CIFAR 100 0.9051 CLEVR Counts 0.2966 CLEVR Distance 0.2124 Country211 0.343981 Describable Textures 0.706383 EuroSAT 0.654815 FGVC Aircraft 0.714055 Food 101 0.956792 GTSRB 0.677514 ImageNet Sketch 0.727308 ImageNet v2 0.773 ImageNet A 0.6988 ImageNet O 0.381 ImageNet R 0.929367 KITTI Vehicle Distance 0.336146 MNIST 0.8579 ObjectNet 0.765156 Oxford Flowers 102 0…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy