⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool ; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3 dolmino mix 100B 1025 Olmo 3 32B: allenai/dolma3 dolmino mix 100B 1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 899M 1.42M TinyMATH PoT Math (synth) 241M 729K CraneMath Math (synth) 5.62B 6.55M MegaMatt Math (synth) 3.88B 6.79M Dolmino Math Math (synth) 10.7B 21M StackEdu (FIM) Code 21.4B 32M CraneCode Python (synth) 18.8B 19.7M Reddit To Flashcards QA (synth) 21.6B 370M Wiki To RCQA QA (synth) 4.22B 22.3M Nemotron Synth QA QA (synth) 487B 972M Math Meta Reasoning Thinking (synth) 1.05B 984K Code Meta Reasoning Thinking (synth) 1.27B 910K Program Verifiable Thinking (synth) 438M 384K OMR Rewrite FullThoughts Thinking (synth) 850M 291K QWQ Reasoning Traces Thinking (synth) 4.77B 438K General Reasoning Mix Thinking (synth) 2.48B 668K Gemini Reasoning Traces Thinking (synth) 246M 55.2K Llama Nemotron Reasoning Traces Thinking (synth) 20…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy