π AutoMathText V2: A 2.46 Trillion Token AI Curated STEM Pretraining Dataset π AutoMathText v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset. π Share your use case here π AutoMathText V2 consists of 2.46 trillion tokens of high quality, deduplicated text spanning web content, mathematics, code, reasoning, and bilingual data. This dataset was meticulously curated using a three tier deduplication pipeline and AI powered quality assessment to provide superior training data for large language models. Our dataset combines 50+ premium data sources with advanced processing techniques, including semantic deduplication , contamination detection , and intelligent text cleaning to deliver exceptional model performance across diverse domains. π― What makes AutoMathText V2 special? π’ STEM Concentration : Specially optimized for STEM content (especially Math) π Triple Deduplication : Exact β Fuzzy (MinHash+LSH) β Semantic (GTE embeddings) π€ AI Quality Assessment : Qwen2 based classifier with multi source score fusion π§Ή Advanced Text Cleaβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy