Ettin Pre training Data Phase 1 of 3 : Diverse pre training data mixture (1.7T tokens) used to train the Ettin model suite. This dataset contains the pre training phase data used to train all Ettin encoder and decoder models. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description : : : : DCLM 837.2 49.1% High quality web crawl data CC Head 356.6 20.9% Common Crawl head documents Starcoder 263.9 15.5% Code repositories and files Reddit 80.3 4.7% Social discussion threads PeS2o 57.3 3.4% Scientific papers Arxiv 28.0 1.6% Academic preprints StackExchange 19.6 1.2% Q&A forums Tulu Flan 16.6 1.0% Instruction following data Open Web Math 12.7 0.7% Mathematical content Algebraic StackExchange 12.6 0.7% Math Q&A CC News 7.3 0.4% News articles Wikipedia 7.3 0.4% Encyclopedia articles Total 1,704.7 100.0% Diverse mixture for foundation training 🚀 Usage For pre training, see the ModernBERT repo: https://github.com/AnswerDotAI/ModernBERT Direct Access 📁 Structure Each folder contains one data source in MDS (Mosaic Data Shard) format: arxiv/ Academic papers from ArXiv books/ Liter…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy