FG CLIP 2: A Bilingual Fine grained Vision language Alignment Model Code: https://github.com/360CVGroup/FG CLIP Project page: https://360cvgroup.github.io/FG CLIP FG CLIP 2 is the foundation model for fine grained vision language understanding in both English and Chinese. Across 29 datasets and 8 diverse tasks, it consistently surpasses recent strong baselines such as SigLIP 2 and MetaCLIP 2, achieving the best reported performance to date in both languages. FG CLIP 2: A Bilingual Fine grained Vision language Alignment Model Chunyu Xie , Bin Wang , Fanjing Kong, Jincheng Li, Dawei Liang, Ji Ao, Dawei Leng†, Yuhui Yin( Equal Contribution, †Corresponding Author) FG CLIP: Fine Grained Visual and Textual Alignment (code branch: v1.0) Chunyu Xie , Bin Wang , Fanjing Kong, Jincheng Li, Dawei Liang, Gengshen Zhang, Dawei Leng†, Yuhui Yin ( Equal Contribution, †Corresponding Author) Quick Start 🤗 Load Model Retrieval Dense feature effect display Citation If you find FG CLIP 2 useful for your research and applications, please cite using this BibTeX: License This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy