VideoRLVR Data · Verifiable video reasoning data for RLVR Training and evaluation data accompanying Video Models Can Reason with Verifiable Rewards . VideoRLVR Data is built for a stricter question: Can a video model generate a full visual trajectory that is not only realistic, but also rule consistent, executable, and automatically verifiable? The dataset contains procedurally generated visual reasoning tasks from three domains: Maze, FlowFree, and Sokoban. Each sample pairs an initial visual state and task prompt with a ground truth solution video. These domains are designed for reinforcement learning with verifiable rewards, where generated videos can be parsed and checked by symbolic rule based evaluators. What's in this repo Path Description train/ Training split for VideoRLVR supervised fine tuning and RL initialization test/ Held out evaluation split generated with disjoint random seeds train/maze/ Maze training videos and metadata train/flowfree/ FlowFree training videos and metadata train/sokoban/ Sokoban training videos and metadata test/maze/ Maze test videos and metadata test/flowfree/ FlowFree test videos and metadata test/sokoban/ Sokoban test videos and metadata trai…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy