CogAgent CogAgent is an open source visual language model improved based on CogVLM . 📖 Paper: https://arxiv.org/abs/2312.08914 🚀 GitHub: For more information such as demo, fine tuning, and query prompts, please refer to Our GitHub Reminder 📍 This is the cogagent vqa version of CogAgent checkpoint. We have open sourced 2 versions of CogAgent checkpoints, and you can choose one based on your needs. 1. cogagent chat : This model has strong capabilities in GUI Agent, visual multi turn dialogue, visual grounding, etc. If you need GUI Agent and Visual Grounding functions, or need to conduct multi turn dialogues with a given image, we recommend using this version of the model. 3. cogagent vqa : This model has stronger capabilities in single turn visual dialogue . If you need to work on VQA benchmarks (such as MMVET, VQAv2), we recommend using this model. Introduction CogAgent 18B has 11 billion visual and 7 billion language parameters. CogAgent demonstrates strong performance in image understanding and GUI agent: 1. CogAgent 18B achieves state of the art generalist performance on 9 cross modal benchmarks , including: VQAv2, MM Vet, POPE, ST VQA, OK VQA, TextVQA, ChartQA, InfoVQA, DocVQ…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy