Dataset Description: The SDG SynHuman is a large scale synthetic video dataset of digital humans rendered in diverse indoor and outdoor 3D environments. The dataset contains 236,937 clips, totaling approximately 5,841 hours of video, and is designed to support training and post training of NVIDIA Cosmos world foundation models and related physical AI research. Each sample is a temporally coherent 60 120 second video clip rendered at 1080p and 30 fps. Clips contain multiple digital humans performing animation sequences in a sampled 3D environment with a controlled camera trajectory. The dataset spans 4,050 digital human assets, 8,184 unique animations, 198 indoor environments, 200 outdoor city environments, and 14 camera motion scenarios, providing broad variation in human appearance, motion, scene context, lighting, and camera behavior. The dataset is intended as a controllable synthetic supplement to real world human video data for applications such as world model pretraining and post training, camera motion generalization, depth and geometry aware learning, human scene interaction modeling, and physical AI research. This dataset is fully synthetic. It contains no real world image…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy