Qwen2.5 Math RM 72B Introduction Qwen2.5 Math RM 72B is specifically designed to guide the Qwen2.5 Math model throughout the training process by offering more granular feedback on the quality of reasoning and intermediate steps, ultimately facilitating more robust model improvements. Key Highlights: Multilingual and Multi Modal Support: Offers preference signals across two languages (Chinese and English) and in dual modes (Chain of Thought and Tool integrated Reasoning), enhancing versatility. Model Training Guide: Training Data Enhancement: Employs a data selection process via reward model scoring combined with Rejection Sampling to incrementally enhance the quality of responses Reinforcement Learning Training: Integrates seamlessly into the reinforcement learning training and provide effective reward signal, further improving model performance. Inference Boosting: Best of N: By leveraging a combination of response sampling and Best of N strategies, we choose the response of top score judged by reward model, yielding better results with spending more inference time. For example, Qwen2.5 Math 1.5B Instruct obtains 83.9 on MATH in RM@8 setting and even surpasses the performance of Q…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy