MMBERT Decay Phase Data Phase 3 of 3 : Annealed language learning decay phase (100B tokens) with massive multilingual expansion to 1833 languages. 📊 Data Composition NOTE: there are multiple decay data mixtures: this mixture described below is the Decay Cont mixture. However, the data in this repository is the Decay Eng. If you are interested in the others, please let me know so I can prioritize it. Data Source Tokens (B) Percentage Description : : : : FineWeb2 78.5 76.0% High quality multilingual web crawl data Wikipedia (MegaWika) 9.5 9.2% Encyclopedia articles (1833 languages) Arxiv 3.3 3.2% Academic preprints Textbooks (ProLong) 3.1 3.0% Educational content Code (ProLong) 2.8 2.7% Code repositories and files Books 2.2 2.1% Literature and reference books DCLM (Dolmino) 2.0 2.0% High quality English web data Tulu Flan 1.0 1.0% Instruction following data Starcoder 0.5 0.5% Code repositories Dolmino Math 0.5 0.5% Mathematical content Total 103.3 100.0% Optimized for rapid language acquisition 🌍 Massive Language Coverage This phase dramatically expands language coverage to 1833 languages , implementing the novel Cascading Annealed Language Learning (ALL) approach: Temperature Sche…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy