SwallowMath v2 Resources 📑 arXiv : Read our paper for detailed methodology at arXiv:2505.02881. 🤗 Sister Dataset : Discover SwallowCode2, our companion dataset for code generation. 🧮 What is it? SwallowMath v2 is a large scale mathematical dataset containing 32 billion tokens , developed as the successor to SwallowMath v1. Building on the success of v1, this release aims to construct a larger scale and more permissively licensed corpus to support open and reproducible research on mathematical reasoning for large language models (LLMs). As in our previous dataset SwallowMath v1, SwallowMath v2 employs an LLM driven rewriting approach —removing boilerplate, restoring missing context, and reformatting solutions into clear, step by step explanations. Additionally, we explored multiple rewriting styles and adopted the two most effective ones—Textbook and Q&A—in the final synthesis stage, yielding higher consistency and reasoning quality. Empirical evaluations demonstrate that models trained with SwallowMath v2 achieve stronger performance on GSM Plus and BBH , surpassing other open mathematical datasets. † On the MATH benchmark, the SwallowMath v2 (Q&A) variant performs slightly belo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy