Qwen2.5 3B GRPO Math GSM8K A 3 billion parameter Qwen2.5 model fine tuned with Group Relative Policy Optimization (GRPO) on the GSM8K grade school math dataset. The aim is to turn the compact 3B model into a lightweight but highly capable step by step math reasoner that runs on a single consumer GPU. Try it
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy