Z Image Turbo Fun Controlnet Union News The new control model with more control blocks and inpaint mode is released. Model Features This ControlNet is added on 6 blocks. The model was trained from scratch for 10,000 steps on a dataset of 1 million high quality images covering both general and human centric content. Training was performed at 1328 resolution using BFloat16 precision, with a batch size of 64, a learning rate of 2e 5, and a text dropout ratio of 0.10. It supports multiple control conditions—including Canny, HED, Depth, Pose and MLSD can be used like a standard ControlNet. You can adjust control context scale for stronger control and better detail preservation. For better stability, we highly recommend using a detailed prompt. The optimal range for control context scale is from 0.65 to 0.80. TODO [ ] Train on more data and for more steps. [ ] Support inpaint mode. Results Pose Output Pose Output Canny Output HED Output Depth Output Inference Go to the VideoX Fun repository for more details. Please clone the VideoX Fun repository and create the required directories: Then download the weights into models/Diffusion Transformer and models/Personalized Model. Then run the fi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy