━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ Unlocking the Reasoning Potential of Language Model From Pretraining to Posttraining ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 🤗 HuggingFace 🤖️ ModelScope 📔 Technical Report Updates [2025.05.30] We scaled the SFT dataset from approximately 500K to 6M instances and continuously expanding the RL training window size from 32K to 48K, the performance of MiMo 7B RL 0530 on AIME24 can be continuously improved and eventually surpass that of DeepSeek R1 (79.8). Benchmark MiMo 7B RL MiMo 7B RL 0530 Mathematics MATH500 (Pass@1) 95.8 97.2 AIME 2024 (Pass@1) 68.2 80.1 AIME 2025 (Pass@1) 55.4 70.2 Code LiveCodeBench v5 (Pass@1) 57.8 60.9 LiveCodeBench v6 (Pass@1) 49.3 52.2 STEM GPQA Diamond (Pass@1) 54.4 60.6 General Alignbench1.1 (Evaluated by GPT4.1) 6.9 7.4 I. Introduction Currently, most successful RL works, including open source research, rely on relatively large base models, e.g., 32B models, particularly for enhancing code reasoning capabilities. Moreover, it was widely considered that achieving uniform and simultaneous improvements in both mathematical and code capabilities within a small model is challenging. Nonetheless…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy