Zebra‑CoT A diverse large scale dataset for interleaved vision‑language reasoning traces. Dataset Description Zebra‑CoT is a diverse large‑scale dataset with 182,384 samples containing logically coherent interleaved text‑image reasoning traces across four major categories: scientific reasoning, 2D visual reasoning, 3D visual reasoning, and visual logic & strategic games. Dataset Structure Each example in Zebra‑CoT consists of: Problem statement : textual description of the question. Problem image : Zero or more images accompanying the problem, depending on its nature. Reasoning image : At least one or more visual aids that support intermediate reasoning steps during problem solving. Text Reasoning Trace : a sequence of text thoughts ( THOUGHT x ) and corresponding visual sketches or diagrams placeholders, such as [problem image 1] and [reasoning image 1] . Final answer : the solution to the problem. Usage To prepare the interleaved text image traces for training, replace all [problem image x] and [reasoning image x] in the text trace with the actual images. We performed careful data cleaning to make sure each image and image placeholder has a one to one mapping. For process supervi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy