AutoMathText-2.5
๐ AutoMathText-2.5: A Foundational High-Quality STEM Training Dataset
๐ AutoMathText-2.5 consists of over 2 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and bilingual data. This dataset was meticulously curated using a three-tier deduplication pipeline and AI-powered quality assessment to provide superior training data for large language models.
Our dataset combines 50+ premium data sources with advanced processing techniques, including semantic deduplication, contamination detection, and intelligent text cleaning to deliver exceptional model performance across diverse domains.
๐ Licensing & Citation
License
Released under AutoMathText Data Agreement for Model Training (See LICENSE).ย
Citation
@misc{automathtext_2_5,
ย title={AutoMathText-2.5: A Foundational High-Quality STEM Training Dataset},
ย author={Zhang, Yifan and Feng, Jichen and Math-AI, Team},
ย year={2026},
ย publisher={Hugging Face},
ย url={https://huggingface.co/datasets/math-ai/AutoMathText-2.5},
}
@article{zhang2025autonomous,
ย title={Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts},
ย author={Zhang, Yifan and Luo, Yifan and Yuan, Yang and Yao, Andrew C},
ย journal={Findings of the Association for Computational Linguistics: ACL 2025},
ย year={2025}
}