mmBERT Mid training Data Phase 2 of 3 : High quality mid training data mixture (600B tokens) with context extension to 8192 tokens. This dataset contains the mid training phase data used to train all mmBERT encoder models. This phase focuses on higher quality data sources and extends the context length from 1024 to 8192 tokens. The data is provided in MDS format ready for use with Composer and the ModernBERT training repository. 📊 Data Composition Data Source Tokens (B) Percentage Description : : : : FineWeb2 506.7 84.3% High quality multilingual web crawl data DCLM (Dolmino) 40.0 6.7% Filtered high quality English web data Starcoder 17.2 2.9% Code repositories and files Arxiv 5.4 0.9% Academic preprints Dolmino Math 4.3 0.7% Mathematical content Books 3.9 0.7% Literature and reference books PeS2o 3.2 0.5% Scientific papers Tulu Flan 3.1 0.5% Instruction following data StackExchange 3.0 0.5% Q&A forums StackExchange (Dolmino) 2.8 0.5% Curated Q&A content Wikipedia (MegaWika) 1.2 0.2% Encyclopedia articles Total 600.8 100.0% High quality data for context extension 🌍 Language Coverage This phase covers 110 languages plus code, with inverse temperature sampling at τ=0.5. Expands fro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy