D2E 480p Project Page · Paper (arXiv) · GitHub · OWA Toolkit Documentation This is the dataset for D2E: Scaling Vision Action Pretraining on Desktop Data for Transfer to Embodied AI . 268.7 hours of synchronized video, audio, and input events from 29 PC games across diverse genres (FPS, open world, sandbox, and more), for training vision action models and game agents. What's included: Video + Audio : H.264 encoded at 480p 60fps with game audio. Fixed 0.5s keyframe intervals and disabled B frame for efficient random seek without sequential decoding. Input events : Keyboard (press/release + key state), mouse (clicks, screen coordinates, raw HID deltas, button state), and active window info—all with nanosecond timestamps synchronized to video frames. OWAMcap format : Built on MCAP (widely adopted in robotics). Indexed for fast random access, crash safe writes, and standardized message schemas that work across different datasets without custom parsing. Recommended for: Training game agents with vision action trajectories, pretraining vision action models for transfer to embodied AI (robotic manipulation, navigation), or world model / video generation training (use D2E Original for FHD/…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy