Pyramid Vision Transformer (tiny sized model) Pyramid Vision Transformer (PVT) model pre trained on ImageNet 1K (1 million images, 1000 classes) at resolution 224x224, and fine tuned on ImageNet 2012 (1 million images, 1,000 classes) at resolution 224x224. It was introduced in the paper Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions by Wenhai Wang, Enze Xie, Xiang Li, Deng Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao and first released in this repository. Disclaimer: The team releasing PVT did not write a model card for this model so this model card has been written by [Rinat S. [@Xrenya]](https://huggingface.co/Xrenya). Model description The Pyramid Vision Transformer (PVT) is a transformer encoder model (BERT like) pretrained on ImageNet 1k (also referred to as ILSVRC2012), a dataset comprising 1 million images and 1,000 classes, also at resolution 224x224. Images are presented to the model as a sequence of variable size patches, which are linearly embedded. Unlike ViT models, PVT is using a progressive shrinking pyramid to reduce computations of large feature maps at each stage. One also adds a [CLS] token to the beg…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy