Being H0: Vision Language Action Pretraining from Large Scale Human Videos We introduce Being H0 , the first dexterous Vision Language Action model pretrained from large scale human videos via explicit hand motion modeling. News [2025 08 02] : We release the Being H0 codebase and pretrained models! Check our Hugging Face Model Hub for more details. π₯π₯π₯ [2025 07 21] : We publish Being H0 ! Check our paper here. πππ Model Checkpoints Download pre trained models from Hugging Face: Model Type Model Name Parameters Description Motion Model Being H0 GRVQ 8K Motion tokenizer VLA Pre trained Being H0 1B 2508 1B Base vision language action model VLA Pre trained Being H0 8B 2508 8B Base vision language action model VLA Pre trained Being H0 14B 2508 14B Base vision language action model VLA Post trained Being H0 8B Align 2508 8B Fine tuned for robot alignment Dataset We have provided the dataset for post training the VLA model. The dataset is available in Hugging Face: Dataset Type Dataset Name Description VLA Post training h0 post train db 2508 Post training dataset for pretrained Being H0 VLA model Setup Clone repository Create environment Install package Download MANO package Visit Mβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy