📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath 3+) and 54B tokens (FineMath 3+ with InfiMM WebMath 3+) of mathematical educational content filtered from CommonCrawl. To curate this dataset, we trained a mathematical content classifier using annotations generated by LLama 3.1 70B Instruct. We used the classifier to retain only the most educational mathematics content, focusing on clear explanations and step by step problem solving rather than advanced academic papers. The Dataset Curation section details the process for creating the dataset. More details in our paper: https://arxiv.org/abs/2502.02737v1. What is being released? The dataset is released in two versions: FineMath 3+ : 34B tokens, 21.4M documents containing mathematical reasoning and problem solving, formatted with Markdown and LaTeX. FineMath 4+ (a subset of FineMath 3+): 9.6B tokens, 6.7M documents of higher quality with detailed explanations. Models trained on this dataset perform better on GSM8k and MATH. We also release a filtered English text only portion of the InfiMM WebMath 40B dataset, classified using the same approach as FineMath: InfiMM WebMath 3+ : 20.5B tokens, 13.9M documents. InfiMM…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy