Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a massive collection of 50 vision language datasets (training sets only) that were used for the fine tuning of the vision language model Idefics2. Load the dataset To load the dataset, install the library datasets with pip install datasets . Then, to download and load the config ai2d for example. Data fields An example of a sample looks as follows: In images , there is a list of images, to be placed before the text. In texts , there is a conversation between a user and an assistant about the images that is represented by a list of turns. Stats about the datasets in The Cauldron Dataset images Q/A pairs tokens General visual question answering VQAv2 82,772 443,757 1,595,929 COCO QA 46,287 78,736 286,982 Visual7W 14,366 69,817 279,268 A OKVQA 16,539 17,056 236,492 TallyQA 98,680 183,986 738,254 OK VQA 8,998 9,009 38,853 HatefulMemes 8,500 8,500 25,500 VQA RAD 313 1,793 8,418 Captioning LNarratives 507,444 507,444 21,328,731 Screen2Words 15,730 15,743 143,103 VSR 2,157 3,354 10,062 OCR, document understanding, text transcription RenderedText 999,000 999,000 27,207,774 DocVQA 10,189 39…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy