PusaV0.5 Training Dataset Code Repository Model Hub Training Toolkit Dataset Pusa Paper FVDM Paper Follow on X Xiaohongshu Dataset Overview This repository contains the pre encoded training dataset used for fine tuning the Pusa V0.5 video generation model. The dataset consists of 52,695 pre encoded latent samples derived from VIDGEN 1M, total size is 785GB, though Pusa V0.5 was trained using only 16,000 of this dataset. Dataset Structure The dataset is organized into two main directories: videos/ : Contains pre encoded video latents in PyTorch tensor format. Atually, the corresponding videos ( .mp4 files) are also provided in videos/ , you may check them out for more details. captions/ : Contains corresponding text embeddings for each video Dataset Details Total Samples : 52,695 video text embedding pairs Source : Randomly sampled from VIDGEN 1M Format : Pre encoded latents (.pt files) ready for training Used in Pusa V0.5 : 16,000 samples from this dataset were used to train the released Pusa V0.5 model Usage Download the Dataset Unzip the Dataset Using with Mochi Full Finetuner This dataset is designed to work seamlessly with the Mochi Full Finetuner repository for training Pusa o…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy