RoboReward Links: Paper · RoboRewardBench Leaderboard RoboReward is a dataset for training and evaluating general purpose vision language reward models for robotics. Each example pairs a task instruction with a real robot rollout video and a discrete end of episode progress reward score in {1,…,5}. RoboReward is built from large scale real robot corpora including Open X Embodiment (OXE) and RoboArena . Because OXE is success heavy, we generate additional negatives and near misses using: Counterfactual relabeling: keep the same rollout video but swap in alternative task instructions that would score lower given the final state. Temporal clipping: truncate successful episodes to produce partial progress outcomes for the original task. RoboReward additionally integrates organic successes and failures from RoboArena. This dataset is used to train the RoboReward reward models (4B/8B) (finetuning Qwen 3 VL) and to evaluate reward accuracy on a standardized benchmark. The test split is fully human verified and released as RoboRewardBench . What’s in this dataset Splits and size Total: 54,135 examples Train: 45,072 Validation: 6,232 Test: 2,831 ( RoboRewardBench ; human verified) Fields Ea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy