Vision and Language Transformer (ViLT), pre trained only Vision and Language Transformer (ViLT) model pre trained on GCC+SBU+COCO+VG (200k steps). It was introduced in the paper ViLT: Vision and Language Transformer Without Convolution or Region Supervision by Kim et al. and first released in this repository. Note: this model only includes the language modeling head. Disclaimer: The team releasing ViLT did not write a model card for this model so this model card has been written by the Hugging Face team. Intended uses & limitations You can use the raw model for masked language modeling given an image and a piece of text with [MASK] tokens. How to use Here is how to use this model in PyTorch: Training data (to do) Training procedure Preprocessing (to do) Pretraining (to do) Evaluation results (to do) BibTeX entry and citation info
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy