Model card for ViT B 16 SigLIP2 Model Details A SigLIP 2 Vision Lanuage model trained on WebLI. This model has been converted for use in OpenCLIP from the original JAX checkpoints in Big Vision. Model Details Model Type: Contrastive Image Text, Zero Shot Image Classification. Original: https://github.com/google research/big vision Dataset: WebLI Papers: SigLIP 2: Multilingual Vision Language Encoders with Improved Semantic Understanding, Localization, and Dense Features: https://arxiv.org/abs/2502.14786 Sigmoid loss for language image pre training: https://arxiv.org/abs/2303.15343 Model Usage Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy