SpatialVID: A Large Scale Video Dataset with Spatial Annotations Jiahao Wang 1 Yufeng Yuan 1 Rujie Zheng 1 Youtian Lin 1 Jian Gao 1 Lin Zhuo Chen 1 Yajie Bao 1 Yi Zhang 1 Chang Zeng 1 Yanxi Zhou 1 Xiaoxiao Long 1 Hao Zhu 1 Zhaoxiang Zhang 2 Xun Cao 1 Yao Yao 1† 1 Nanjing University 2 Institute of Automation, Chinese Academy of Science Equal Contribution †Corresponding Author CVPR 2026 SpatialVID HQ Directory Structure Dataset Download You can download the entire SpatialVID HQ dataset using the following command: The whole dataset is approximately 3.53TB in size. We have split the dataset into 74 groups for easier management. Each group contains approximately 14GB of video data and 1.5GB of annotation data, with naming conventions following the format group 0 (e.g., group 0001 , group 0002 ). For downloading specific files (instead of the full dataset), please refer to the download SpatialVID.py provided in our GitHub repository. Usage Guide 1. Unzipping Group Files After downloading the group files (in .tar.gz format), use the tar command to extract their contents. For example: 2. Using the Metadata File The SpatialVID HQ metadata.csv file contains comprehensive metadata for all vi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy