Wan2.2 Fun Reward LoRAs Introduction We explore the Reward Backpropagation technique 1 2 to optimized the generated videos by Wan2.2 Fun for better alignment with human preferences. We provide the following pre trained models (i.e. LoRAs) along with the training script. You can use these LoRAs to enhance the corresponding base model as a plug in or train your own reward LoRA. For more details, please refer to our GitHub repo. Name Base Model Reward Model Hugging Face Description Wan2.2 Fun A14B InP high noise HPS2.1.safetensors Wan2.2 Fun A14B InP (high noise) HPS v2.1 🤗Link Official HPS v2.1 reward LoRA ( rank=128 and network alpha=64 ) for Wan2.2 Fun A14B InP (high noise). It is trained with a batch size of 8 for 5,000 steps. Wan2.2 Fun A14B InP low noise HPS2.1.safetensors Wan2.2 Fun A14B InP (low noise) MPS 🤗Link Official HPS v2.1 reward LoRA ( rank=128 and network alpha=64 ) for Wan2.2 Fun A14B InP (low noise). It is trained with a batch size of 8 for 2,700 steps. Wan2.2 Fun A14B InP high noise MPS.safetensors Wan2.2 Fun A14B InP (high noise) HPS v2.1 🤗Link Official MPS reward LoRA ( rank=128 and network alpha=64 ) for Wan2.2 Fun A14B InP (high noise). It is trained with a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy