REVISOR 25k A multi task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k samples spanning Video QA and Temporal Grounding tasks. Dataset Structure The dataset is organized into 4 subsets: Subset Task Samples Description video r1 Video QA 20,855 Multiple choice video question answering time r1 Temporal Grounding 2,500 Locate time intervals in videos cg bench Temporal Grounding 1,167 CG Bench temporal grounding rextime Temporal Grounding 837 ReXTime temporal grounding Total: 25,359 samples Data Fields Field Type Description video string Relative path to the video file video filename string Video filename sample fps int Sampling FPS used for training original fps string Original video FPS duration string Video duration in seconds video id string Unique video identifier conversations string (JSON) Full conversation (system + user messages) system prompt string System prompt with tool definitions query string Raw user query ground truth string Ground truth answer (letter or time interval) data source string Source dataset identifier env name string Tool environment name ability string Task type: "video qa" or "temporal gr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy