PhotoMaker V2 Model Card Project Page Paper (ArXiv) Code 🤗 Gradio demo Introduction Users can input one or a few face photos, along with a text prompt, to receive a customized photo or painting within seconds (no training required!). Additionally, this model can be adapted to any base model based on SDXL or used in conjunction with other LoRA modules. Realistic results Stylization results More results can be found in our project page Model Details It mainly contains two parts corresponding to two keys in loaded state dict: 1. id encoder includes finetuned OpenCLIP ViT H 14 and a few fuse layers. 2. lora weights applies to all attention layers in the UNet, and the rank is set to 64. Usage You can directly download the model in this repository. You also can download the model in python script: Then, please follow the instructions in our GitHub repository. Limitations The model's customization performance degrades on Asian male faces. The model still struggles with accurately rendering human hands. Bias While the capabilities of image generation models are impressive, they can also reinforce or exacerbate social biases. Citation BibTeX:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy