VisCoT Dataset Card There is a shortage of multimodal datasets for training multi modal large language models (MLLMs) that require to identify specific regions in an image for additional attention to improve response performance. This type of dataset with grounding bbox annotations could possibly help the MLLM output intermediate interpretable attention area and enhance performance. To fill the gap, we curate a visual CoT dataset. This dataset specifically focuses on identifying critical regions within images, a feature essential for models to concentrate on relevant visual elements to improve response accuracy. Each data sample consists of a question, answer, and a corresponding visual bounding box across five domains. Some data samples also include extra detailed reasoning steps. To ensure a robust foundation for detailed visual and textual analysis, our dataset deliberately integrates a diverse selection of data including text/doc, fine grained understanding, charts, general VQA, and relation reasoning . These data domains are deliberately chosen to cultivate a comprehensive skill set across varied analytical tasks: 1) Text/doc enhances MLLM's capabilities on OCR and contextual…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy