Vchitect T2V Dataverse Vchitect Team 1   1 Shanghai Artificial Intelligence Laboratory  Paper Project Page Data Overview The Vchitect T2V Dataverse is the core dataset used to train our text to video diffusion model, Vchitect 2.0: Parallel Transformer for Scaling Up Video Diffusion Models. It comprises 14 million high quality videos collected from the Internet, each paired with detailed textual captions. This large scale dataset enables the model to learn rich video text alignments and generate temporally coherent video content from textual prompts. For more technical details, data processing procedures, and model training strategies, please refer to our paper. BibTex Disclaimer We disclaim responsibility for user generated content. The model was not trained to realistically represent people or events, so using it to generate such content is beyond the model's capabilities. It is prohibited for pornographic, violent and bloody content generation, and to generate content that is demeaning or harmful to people or their environment, culture, religion, etc. Users are solely liable for their actions. The project contributors are not legally affiliated with, nor accountable for…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy