Stable Diffusion v2 base Model Card This model card focuses on the model associated with the Stable Diffusion v2 base model, available here. The model is trained from scratch 550k steps at resolution 256x256 on a subset of LAION 5B filtered for explicit pornographic material, using the LAION NSFW classifier with punsafe=0.1 and an aesthetic score = 4.5 . Then it is further trained for 850k steps at resolution 512x512 on the same dataset on images with resolution = 512x512 . Use it with the stablediffusion repository: download the 512 base ema.ckpt here. Use it with 🧨 diffusers Model Details Developed by: Robin Rombach, Patrick Esser Model type: Diffusion based text to image generation model Language(s): English License: CreativeML Open RAIL++ M License Model Description: This is a model that can be used to generate and modify images based on text prompts. It is a Latent Diffusion Model that uses a fixed, pretrained text encoder (OpenCLIP ViT/H). Resources for more information: GitHub Repository. Cite as: @InProceedings{Rombach 2022 CVPR, author = {Rombach, Robin and Blattmann, Andreas and Lorenz, Dominik and Esser, Patrick and Ommer, Bj\"orn}, title = {High Resolution Image Synthe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy