Donut (base sized model, pre trained only) Donut model pre trained only. It was introduced in the paper OCR free Document Understanding Transformer by Geewok et al. and first released in this repository. Disclaimer: The team releasing Donut did not write a model card for this model so this model card has been written by the Hugging Face team. Model description Donut consists of a vision encoder (Swin Transformer) and a text decoder (BART). Given an image, the encoder first encodes the image into a tensor of embeddings (of shape batch size, seq len, hidden size), after which the decoder autoregressively generates text, conditioned on the encoding of the encoder. Intended uses & limitations This model is meant to be fine tuned on a downstream task, like document image classification or document parsing. See the model hub to look for fine tuned versions on a task that interests you. How to use We refer to the documentation which includes code examples. BibTeX entry and citation info
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy