Model card for ViT SO400M 14 SigLIP A SigLIP (Sigmoid loss for Language Image Pre training) model trained on WebLI. This model has been converted to PyTorch from the original JAX checkpoints in Big Vision. These weights are usable in both OpenCLIP (image + text) and timm (image only). Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Original: https://github.com/google research/big vision Dataset: WebLI Papers: Sigmoid loss for language image pre training: https://arxiv.org/abs/2303.15343 Model Usage With OpenCLIP With timm (for image embeddings) Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy