WorldBench WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models and vision language models. There are two components: Video based: This is a benchmark designed to evaluate video to video world foundation models such as Cosmos. It consists of videos 132 frames long of 425 simulated scenes with RGB, Normals, Depth, Flow, and Segmentations. Text based: This is a subset which adds text based questions to 181 videos from the video benchmark. Questions can be both multiple choice or binary. The video based benchmark can be found in /scenes. There are 4 high level categories for different physics concepts being tested. Within each, there are 3 5 scenes each with 25 50 variations. The text based benchmark is in /textual questions. There are 4 JSON files, one per category. Code to run the evaluation for this benchmark along with instructions can be found here: https://drive.google.com/file/d/1TNHfV mKiidl1 eFWJyctBOodWJnCajA/view?usp=sharing
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy