CogVideoX 5B I2V ๐ Read in English ๐ค Huggingface Space ๐ Github ๐ arxiv ๐ Visit Qingying and API Platform for the commercial version of the video generation model Model Introduction CogVideoX is an open source video generation model originating from Qingying. The table below presents information related to the video generation models we offer in this version. Model Name CogVideoX 2B CogVideoX 5B CogVideoX 5B I2V (This Repository) Model Description Entry level model, balancing compatibility. Low cost for running and secondary development. Larger model with higher video generation quality and better visual effects. CogVideoX 5B image to video version. Inference Precision FP16 (recommended) , BF16, FP32, FP8 , INT8, not supported: INT4 BF16 (recommended) , FP16, FP32, FP8 , INT8, not supported: INT4 Single GPU Memory Usage SAT FP16: 18GB diffusers FP16: from 4GB diffusers INT8 (torchao): from 3.6GB SAT BF16: 26GB diffusers BF16: from 5GB diffusers INT8 (torchao): from 4.4GB Multi GPU Inference Memory Usage FP16: 10GB using diffusers BF16: 15GB using diffusers Inference Speed (Step = 50, FP/BF16) Single A100: ~90 seconds Single H100: ~45 seconds Single A100: ~180 seconds Single H1โฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy