๐งโโ๏ธ Qwen3 4B RPG Roleplay V2 (GRPO) Aligning Characters with Deeper Personas A new version trained with GRPO for more consistent, high quality, and aligned character roleplaying. ๐ Model Overview Welcome to V2! I'm Chun (@chun121), and this is the next evolution of the Qwen3 4B Roleplay model. This version moves beyond standard fine tuning and leverages GRPO (Generative Responsive Preference Optimization) to align the model's behavior with the core principles of great roleplaying. ๐ญ ๐ฌ ๐ง โ๏ธ Character Consistency High Quality Dialogue Intent Understanding Structured Format Maintains strong persona adherence Detailed, engaging non generic responses Comprehends user questions & scenarios Uses <thinking> analysis process Built on the unsloth/Qwen3 4B Base , this LoRA was trained not just to predict text, but to generate responses that are actively rewarded for being in character, high quality, and contextually aware. It's designed for creators who need AI characters that are not only conversational but also consistent and deeply aligned with their defined personas. ๐ Technical Specifications ๐ง Feature ๐ Details Base Model unsloth/Qwen3 4B Base Architecture Transformer LLMโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy