🐱 PixArt Σ Model Card Model PixArt Σ consists of pure transformer blocks for latent diffusion: It can directly generate 1024px, 2K and 4K images from text prompts within a single sampling process. Source code is available at https://github.com/PixArt alpha/PixArt sigma. Model Description Developed by: PixArt Σ Model type: Diffusion Transformer based text to image generative model License: CreativeML Open RAIL++ M License Model Description: This is a model that can be used to generate and modify images based on text prompts. It is a Transformer Latent Diffusion Model that uses one fixed, pretrained text encoders (T5) and one latent feature encoder (VAE). Resources for more information: Check out our GitHub Repository and the PixArt Σ report on arXiv. Model Sources For research purposes, we recommend our generative models Github repository (https://github.com/PixArt alpha/PixArt sigma), which is more suitable for both training and inference and for which most advanced diffusion sampler like SA Solver will be added over time. Hugging Face provides free PixArt Σ inference. Repository: https://github.com/PixArt alpha/PixArt sigma Demo: https://huggingface.co…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy