FG CLIP: Fine Grained Visual and Textual Alignment FG CLIP: Fine Grained Visual and Textual Alignment Chunyu Xie , Bin Wang , Fanjing Kong, Jincheng Li, Dawei Liang, Gengshen Zhang, Dawei Leng†, Yuhui Yin( Equal Contribution, ✝Corresponding Author) Model Framework FG CLIP’s training proceeds in two stages: the first stage leverages global level caption image pairs to achieve initial fine grained alignment, while the second stage supplements these with additional region level captions, including detailed region captions and positive/negative region descriptions to further refine the alignment. Quick Start 🤗 Load Model Retrieval Dense feature effect display Citation If you find FG CLIP useful for your research and applications, please cite using this BibTeX: License This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses. The content of this project itself is licensed under the Apache license 2.0.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy