Model Card for StreetCLIP StreetCLIP is a robust foundation model for open domain image geolocalization and other geographic and climate related tasks. Trained on an original dataset of 1.1 million street level urban and rural geo tagged images, it achieves state of the art performance on multiple open domain image geolocalization benchmarks in zero shot, outperforming supervised models trained on millions of images. Model Description StreetCLIP is a model pretrained by deriving image captions synthetically from image class labels using a domain specific caption template. This allows StreetCLIP to transfer its generalized zero shot learning capabilities to a specific domain (i.e. the domain of image geolocalization). StreetCLIP builds on the OpenAI's pretrained large version of CLIP ViT, using 14x14 pixel patches and images with a 336 pixel side length. Model Details Model type: CLIP Language: English License: Create Commons Attribution Non Commercial 4.0 Trained from model: openai/clip vit large patch14 336 Model Sources Paper: Preprint Cite preprint as: Uses StreetCLIP has a deep understanding of the visual features found in street level urban and rural scenes and knows how to re…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy