Convolutional Vision Transformer (CvT) CvT 13 model pre trained on ImageNet 1k at resolution 224x224. It was introduced in the paper CvT: Introducing Convolutions to Vision Transformers by Wu et al. and first released in this repository. Disclaimer: The team releasing CvT did not write a model card for this model so this model card has been written by the Hugging Face team. Usage Here is how to use this model to classify an image of the COCO 2017 dataset into one of the 1,000 ImageNet classes:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy