A CLIP (Contrastive Language Image Pre training) model trained on DFN 2B. Data Filtering Networks (DFNs) are small networks used to automatically filter large pools of uncurated data. This model was trained on 2B images that were filtered from a pool of 12.8B uncurated image text pairs (12.8B image text pairs from CommonPool 12.8B). This model has been converted to PyTorch from the original JAX checkpoints from Axlearn (https://github.com/apple/axlearn). These weights are directly usable in OpenCLIP (image + text). Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Dataset: DFN 2b Papers: Data Filtering Networks: https://arxiv.org/abs/2309.17425 Examples Seen: 12.8B Model Metrics Eval Dataset Metric : : ImageNet 1k 0.81396 Caltech 101 0.953141 CIFAR 10 0.9836 CIFAR 100 0.8835 CLEVR Counts 0.3338 CLEVR Distance 0.248733 Country211 0.28237 Describable Textures 0.66117 EuroSAT 0.646296 FGVC Aircraft 0.395945 Food 101 0.945861 GTSRB 0.616152 ImageNet Sketch 0.683311 ImageNet v2 0.7453 ImageNet A 0.6676 ImageNet O 0.3915 ImageNet R 0.900033 KITTI Vehicle Distance 0.201125 MNIST 0.8468 ObjectNet 0.739367 Oxford Flowers 102 0.865822 Oxford IIIT Pet 0.954941…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy