UltraData Math 🤗 Dataset 💻 Source Code 🇨🇳 中文 README UltraData Math is a large scale, high quality mathematical pre training dataset totaling 290B+ tokens across three progressive tiers— L1 (170.5B tokens web corpus), L2 (33.7B tokens quality selected), and L3 (88B tokens multi format refined)—designed to systematically enhance mathematical reasoning in LLMs. It has been applied to the mathematical pre training of the MiniCPM Series models. It was introduced in the paper Data Science and Technology Towards AGI Part I: Tiered Data Management. 🆕 What's New [2026.02.09] : UltraData Math , a large scale high quality mathematical pre training dataset with 290B+ tokens across three progressive tiers (L1/L2 preview/L3), is now available on Hugging Face. Released as part of the UltraData ecosystem. 🔥🔥🔥 [2026.02.10] : UltraData Math tops the Hugging Face Datasets Trending list, reaching the 1 spot! ⭐️⭐️⭐️ 📚 Introduction High quality pre training data is crucial for enhancing the mathematical reasoning capabilities of large language models (LLMs). However, existing mathematical pre training data construction schemes have the following shortcomings: HTML Parsing : General parsers (suc…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy