Model Card: GroupViT This checkpoint is uploaded by Jiarui Xu. Model Details The GroupViT model was proposed in GroupViT: Semantic Segmentation Emerges from Text Supervision by Jiarui Xu, Shalini De Mello, Sifei Liu, Wonmin Byeon, Thomas Breuel, Jan Kautz, Xiaolong Wang. Inspired by CLIP, GroupViT is a vision language model that can perform zero shot semantic segmentation on any given vocabulary categories. Model Date June 2022 Abstract Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end to end deep learning systems, grouping of image regions usually happens implicitly via top down supervision from pixel level recognition labels. Instead, in this paper, we propose to bring back the grouping mechanism into deep networks, which allows semantic segments to emerge automatically with only text supervision. We propose a hierarchical Grouping Vision Transformer (GroupViT), which goes beyond the regular grid structure representation and learns to group image regions into progressively larger arbitrary shaped segments. We train GroupViT jointly with a text encoder on a large scale image text dataset…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy