EgoScalerV2 Dataset This dataset accompanies our work on Developing Vision Language Action Model from Egocentric Videos . It provides 6DoF object trajectories paired with egocentric visual observations and natural language action descriptions, formatted in the LeRobot v2.0 schema so it can be consumed directly by LeRobot compatible pipelines. 🌐 Project page: https://biscue5.github.io/egovla project page/ 📄 Paper: Developing Vision Language Action Model from Egocentric Videos (arXiv:2509.21986) 🧰 Format: LeRobot v2.0 (Parquet + MP4) 🪪 License: Apache 2.0 Dataset Structure meta/info.json: Citation If you use this dataset, please cite: The data construction pipeline builds on:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy