Kosmos 2: Grounding Multimodal Large Language Models to the World [An image of a snowman warming himself by a fire.] This Hub repository contains a HuggingFace's transformers implementation of the original Kosmos 2 model from Microsoft. How to Get Started with the Model Use the code below to get started with the model. Tasks This model is capable of performing different tasks through changing the prompts. First, let's define a function to run a prompt. Click to expand Here are the tasks Kosmos 2 could perform: Click to expand Multimodal Grounding • Phrase Grounding • Referring Expression Comprehension Multimodal Referring • Referring expression generation Perception Language Tasks • Grounded VQA • Grounded VQA with multimodal referring via bounding boxes Grounded Image captioning • Brief • Detailed Draw the bounding bboxes of the entities on the image Once you have the entities , you can use the following helper function to draw their bounding bboxes on the image: Click to expand Here is the annotated image: BibTex and citation info
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy