VisionQ 1k v4.4 A human verified benchmark of per region visual quality annotations on qualitative comparison figures from computer vision papers. 1. Overview VisionQ 1k v4.4 is a human verified benchmark designed to train and evaluate vision language models (VLMs) on per region visual quality judgement in academic figures. The dataset spans 1,399 source papers drawn from CVPR / ICCV / ECCV / NeurIPS / SIGGRAPH 2023–2024 and contains 3,354 qualitative candidate figures with 48,167 annotated bounding boxes . Of the 1,399 papers, 908 are approved (passed figure level and group level human review); these approved papers carry 4,365 group level qualitative claim records and document 857 distinct proposed methods identified across the corpus. Each bounding box is annotated with a 10 field record encoding geometry, method identity, scene/row position, and semantic role. Source PDFs are not redistributed; only extracted figure crops and derived metadata are released. The full corpus is released; benchmark claims are computed over the approved subset (908 papers / 4,365 group claims) unless otherwise stated. 2. Motivation Existing visual quality benchmarks operate at image level (a single…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy