Model card for Pix2Struct Pretrained weights This model is the pretrained version of Pix2Struct , use this model for fine tuning purposes only. Table of Contents 0. TL;DR 1. Using the model 2. Contribution 3. Citation TL;DR Pix2Struct is an image encoder text decoder model that is trained on image text pairs for various tasks, including image captionning and visual question answering. The full list of available models can be found on the Table 1 of the paper: The abstract of the model states that: Visually situated language is ubiquitous—sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domainspecific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image to text model for purely visual language understanding, which can be finetuned on tasks containing visually situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provide…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy