QuiltNet B 32 Description QuiltNet B 32 is a CLIP ViT B/32 vision language foundation model trained on the Quilt 1M dataset curated from representative histopathology videos. It can perform various vision language processing (VLP) tasks such as cross modal retrieval, image classification, and visual question answering. QuiltNet establishes new state of the art in a wide range of standard datasets, and substantially outperforms prior VLP approaches: Citation Uses As per the original OpenAI CLIP model card, this model is intended as a research output for research communities. We hope that this model will enable researchers to better understand and explore zero shot, arbitrary image classification. We also hope it can be used for interdisciplinary studies of the potential impact of such model. The OpenAI CLIP paper includes a discussion of potential downstream impacts to provide an example for this sort of analysis. Direct Use Zero shot image classification, image and text retrieval, among others. Downstream Use Image classification and other image task fine tuning, linear probe image classification, image generation guiding and conditioning, among others. Intended Use The model is in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy