Rethinking Video Generation Model for the Embodied World If you like our project, please give us a star ⭐ on GitHub for the latest update. Key features 4M robotic video clips(10K+ hours) for large scale video generation training. 1300+ fine grained robotic skills , covering diverse actions and task primitives. Multi modal physical annotations , including RGB, depth, and optical flow . Multi robot and multi task diversity , spanning various robot types, scenarios, and action skills. Rich object interactions , enabling complex and realistic robot behavior modeling. Dataset Structure RoVid X provides structured annotations for each video clip in JSON format, where each entry is indexed by the video filename. Download You can download RoVid X directly from Hugging Face using the official CLI. 📚 Citation If you find this dataset useful, please cite our paper: bibtex @article{deng2026rethinking, title={Rethinking Video Generation Model for the Embodied World}, author={Deng, Yufan and Pan, Zilin and Zhang, Hongyu and Li, Xiaojie and Hu, Ruoqing and Ding, Yufei and Zou, Yiming and Zeng, Yan and Zhou, Daquan}, journal={arXiv preprint arXiv:2601.15282}, year={2026} }
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy