Qwen2.5 Math PRM 7B Introduction In addition to the mathematical Outcome Reward Model (ORM) Qwen2.5 Math RM 72B, we release the Process Reward Model (PRM), namely Qwen2.5 Math PRM 7B and Qwen2.5 Math PRM 72B. PRMs emerge as a promising approach for process supervision in mathematical reasoning of Large Language Models (LLMs), aiming to identify and mitigate intermediate errors in the reasoning processes. Our trained PRMs exhibit both impressive performance in the Best of N (BoN) evaluation and stronger error identification performance in ProcessBench. Model Details For more details, please refer to our paper. Requirements transformers =4.40.0 for Qwen2.5 Math models. The latest version is recommended. [!Warning] 🚨 This is a must because transformers integrated Qwen2.5 codes since 4.37.0 . For requirements on GPU memory and the respective throughput, see similar results of Qwen2 here. Quick Start [!Important] Qwen2.5 Math PRM 7B is a process reward model typically used for offering feedback on the quality of reasoning and intermediate steps rather than generation. Prerequisites Step Separation: We recommend using double line breaks ("\n\n") to separate individual steps within the s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy