X WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboTwin 2.0 fine tuning dataset used to train the X WAM unified 4D World Action Model. It packages dual arm bimanual manipulation demonstrations into a unified multi view RGB D video + low dimensional state/action format, where each episode provides synchronized RGB videos, depth videos, dual arm end effector proprioception, actions, and a large pool of paraphrased language instructions. Source benchmark : RoboTwin 2.0 Embodiment : Dual arm (bimanual) manipulator Tasks : 50 dual arm manipulation tasks Modalities : 3 camera views × (RGB + Depth) + dual arm EE proprioception + dual arm EE actions + language : : Episodes 27,500 Total frames 6,138,940 Avg. frames / episode ~223 Camera views 3 (1 head + 2 wrist) Video resolution 320 × 240, H.264 Frame rate ~16.7 fps Instructions / episode ~120–130 (paraphrase augmentations) Total size ~94 GB Dataset Structure metadata.json maps each episode key to its number of frames, e.g.: Camera Views View Type Description : : : head camera static Head mounted third person view left camera dynamic Left wrist mounted view right camera dyna…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy