RIG bench Anonymous submission to the NeurIPS 2026 Evaluations & Datasets (E&D) Track . A benchmark for reasoning driven image generation : given visual context (images + instruction + optional demonstration pairs), the model must produce the answer as a single image. 2,000 samples 4 task families × 11 subtasks ~1.4 GB Files Data Format Each line in samples.jsonl corresponds to one benchmark sample. A sample includes the input context, the expected visual answer, task labels. Field Meaning input input images, and optional example images. output Target ground truth answer image. main family Four cognitively demanding domains. subtask Eleven fine grained subtasks Loading
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy