Hunyuan DiT : A Powerful Multi Resolution Diffusion Transformer with Fine Grained Chinese Understanding 混元 DiT: 具有细粒度中文理解的多分辨率Diffusion Transformer [[Arxiv]](https://arxiv.org/abs/2405.08748) [[project page]](https://dit.hunyuan.tencent.com/) [[github]](https://github.com/Tencent/HunyuanDiT) This repo contains the distilled Hunyuan DiT in 🤗 Diffusers format. It supports 25 step text to image generation. Dependency Please install PyTorch first, following the instruction in https://pytorch.org Install the latest version of transformers with pip : Then install the latest github version of 🤗 Diffusers with pip : Example Usage with 🤗 Diffusers 📈 Comparisons In order to comprehensively compare the generation capabilities of HunyuanDiT and other models, we constructed a 4 dimensional test set, including Text Image Consistency, Excluding AI Artifacts, Subject Clarity, Aesthetic. More than 50 professional evaluators performs the evaluation. Model Open Source Text Image Consistency (%) Excluding AI Artifacts (%) Subject Clarity (%) Aesthetics (%) Overall (%) SDXL ✔ 64.3 60.6 91.1 76.3 42.7 PixArt α ✔ 68.3 60.9 93.2 77.5 45.5 Playground 2.5 ✔ 71.9 70.8 94.9 83.3 54.3 SD 3 & 10008 77.1 69.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy