MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We introduce MegaMath, an open math pretraining dataset curated from diverse, math focused sources, with over 300B tokens. MegaMath is curated via the following three efforts: Revisiting web data : We re extracted mathematical documents from Common Crawl with math oriented HTML optimizations, fasttext based filtering and deduplication, all for acquiring higher quality data on the Internet. Recalling Math related code data : We identified high quality math related code from large code training corpus, Stack V2, further enhancing data diversity. Exploring Synthetic data : We synthesized QA style text, math related code, and interleaved text code blocks from web data or code data. MegaMath Compared to Existing Datasets MegaMath is the largest open math pre training dataset to date, surpassing DeepSeekMath (120B) tokens. MegaMath Delivers with High Quality During development, we use extensive experiments to find optimal practice for text extraction, deduplication, fasttext training, etc. Training MegaMath data shows better performance than existing open datasets. Training MegaMath on La…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy