SigLIP HD SigLIP HD is a vision encoder fine tuned from SigLIP 2 So400m/16 512px with fine to coarse supervision . Paper: SigLIP HD by Fine to Coarse Supervision Code: https://github.com/LiheYoung/SigLIP HD SigLIP HD exhibits better performance than SigLIP 2 in MLLMs, especially for OCR scenarios. This repository contains only the vision encoder (no text tower). It is a drop in replacement for the SigLIP 2 vision tower: identical architecture and I/O. To use it in an MLLM, keep your existing SigLIP 2 pipeline and only change the vision tower path to this checkpoint. Usage Citation Acknowledgement This work is built upon SigLIP 2. We sincerely thank the authors for open sourcing their excellent vision encoder.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy