A CLIP (Contrastive Language Image Pre training) model trained on DFN 2B. Data Filtering Networks (DFNs) are small networks used to automatically filter large pools of uncurated data. This model was trained on 2B images that were filtered from a pool of 12.8B uncurated image text pairs (12.8B image text pairs from CommonPool 12.8B). These weights are directly usable in OpenCLIP (image + text). Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Dataset: DFN 2b Papers: Data Filtering Networks: https://arxiv.org/abs/2309.17425 Examples Seen: 12.8B Model Metrics Dataset Metric : : ImageNet 1k 0.76236 Caltech 101 0.942894 CIFAR 10 0.9672 CIFAR 100 0.8347 CLEVR Counts 0.232333 CLEVR Distance 0.245267 Country211 0.19545 Describable Textures 0.575532 EuroSAT 0.54 FGVC Aircraft 0.248503 Food 101 0.91303 GTSRB 0.469913 ImageNet Sketch 0.620684 ImageNet v2 0.682 ImageNet A 0.482133 ImageNet O 0.493 ImageNet R 0.830967 KITTI Vehicle Distance 0.192686 MNIST 0.782 ObjectNet 0.631851 Oxford Flowers 102 0.819895 Oxford IIIT Pet 0.936907 Pascal VOC 2007 0.788528 PatchCamelyon 0.521545 Rendered SST2 0.486546 RESISC45 0.61381 Stanford Cars 0.90735 STL 10 0.97525 SUN397…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy