IDEFICS How do I pronounce the model's name? Watch a Youtube tutorial IDEFICS ( I mage aware D ecoder E nhanced à la F lamingo with I nterleaved C ross attention S ) is an open access reproduction of Flamingo, a closed source visual language model developed by Deepmind. Like GPT 4, the multimodal model accepts arbitrary sequences of image and text inputs and produces text outputs. IDEFICS is built solely on publicly available data and models. The model can answer questions about images, describe visual contents, create stories grounded on multiple images, or simply behave as a pure language model without visual inputs. IDEFICS is on par with the original closed source model on various image text benchmarks, including visual question answering (open ended and multiple choice), image captioning, and image classification when evaluated with in context few shot learning. It comes into two variants: a large 80 billion parameters version and a 9 billion parameters version. We also fine tune the base models on a mixture of supervised and instruction fine tuning datasets, which boosts the downstream performance while making the models more usable in conversational settings: idefics 80b ins…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy