Dataset Card for Docmatix Dataset description Docmatix is part of the Idefics3 release (stay tuned). It is a massive dataset for Document Visual Question Answering that was used for the fine tuning of the vision language model Idefics3. Load the dataset To load the dataset, install the library datasets with pip install datasets . Then, If you want the dataset to link to the pdf files as binaries instead of the images, do: Data fields An example of a sample looks as follows: In images , there is a list of up to 4 images, to be placed before the text. In texts , there is a conversation between a user and an assistant about the images that is represented by a list of turns. Comparison to other DocVQA datasets Dataset images Q/A pairs tokens Document visual question answering Docmatix 2,444,750 9,500,000 390,000,000 DocVQA 10,189 39,463 337,829 TextCaps 21,953 21,953 389,658 TextVQA 21,953 34,602 181,918 ST VQA 17,247 23,121 127,846 OCR VQA 165,746 801,579 6,073,824 VisualMRC 3,027 11,988 168,828 IAM 5,663 5,663 144,216 InfoVQA 2,118 10,074 61,048 Diagram image to text 300 300 22,196 Citation BibTeX:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy