VITRA 1M: Human Hand V L A Dataset Dataset Summary VITRA 1M is a large scale Human Hand Visual Language Action (V L A) dataset constructed as described in the paper Scalable Vision Language Action Model Pretraining for Robotic Manipulation with Real Life Human Activity Videos. It contains 1.2 million short episodes with segmented language annotations, camera parameters (corrected intrinsics/extrinsics), and 3D hand reconstructions (left and right hands) based on the MANO hand model. Each episode is stored as a single .npy metadata file. Project page: https://microsoft.github.io/VITRA/ Note: Current metadata has been manually inspected with an estimated annotation accuracy of around 90%. Future versions will improve metadata quality. Dataset Contents & Size Annotation folder: {dataset name}.tar.gz in root/ . Statistics folder: statistics/{dataset name} angle statistics.json contains dataset statistics. Intrinsics folder: intrinsics/{dataset name} contains the intrinsics of videos in Ego4d and Egoexo4d. Episode counts per dataset: Dataset Number of episodes ego4d cooking and cleaning 454,244 ego4d other 494,439 epic 154,464 egoexo4d 67,053 ssv2 52,718 Extraction instructions: After e…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy