[!WARNING] Idefics2 will NOT work with Transformers version between 4.41.0 and 4.43.3 included. See the issue https://github.com/huggingface/transformers/issues/32271 and the fix https://github.com/huggingface/transformers/pull/32275 [!IMPORTANT] As of April 18th, 2024, Idefics2 is part of the 4.40.0 Transformers pypi release. Please upgrade your Transformers version ( pip install transformers upgrade ). Idefics2 Idefics2 is an open multimodal model that accepts arbitrary sequences of image and text inputs and produces text outputs. The model can answer questions about images, describe visual content, create stories grounded on multiple images, or simply behave as a pure language model without visual inputs. It improves upon Idefics1, significantly enhancing capabilities around OCR, document understanding and visual reasoning. We release under the Apache 2.0 license 2 checkpoints: idefics2 8b base: the base model idefics2 8b: the base model fine tuned on a mixture of supervised and instruction datasets (text only and multimodal datasets) idefics2 8b chatty: idefics2 8b further fine tuned on long conversation Model Summary Developed by: Hugging Face Model type: Multi modal model (im…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy