MotionHub MotionHub is a curated multi domain human motion dataset collection released for training and evaluating generalist motion models. The released version contains motion, language, music, speech, and two person interaction supervision in a unified MotionHub annotation format. This public release is used by VersatileMotion (ECCV 2026) . Every subset listed below has been visually inspected, converted to the repository SMPL H convention, re split where needed, and uploaded after data quality review. Release links: Hugging Face dataset · GitHub repository and preview assets Highlights Scale Language Supervision Audio, Music, and Interaction 20 released subsets 1.11M clips 1,528 h motion 164.98M frames 3.31M text to motion prompts 3.31M motion to text references macro / meso / micro caption levels 5.6K music to dance pairs 44.1K speech/audio to gesture pairs 44.1K script to gesture scripts 23.6K interaction text to motion pairs Modality Previews The previews below are rendered with Three.js from the released SMPL H motion files and grouped by modality rather than task direction. Music and speech examples include the paired source audio where available. Text and Motion Music and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy