Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high quality pool of data considered for the second stage of Olmo 3 32B. Dataset Sources Source Category TinyMATH Mind Math (synth) TinyMATH PoT Math (synth) CraneMath Math (synth) MegaMatt Math (synth) Dolmino Math Math (synth) StackEdu (FIM) Code CraneCode Python (synth) Reddit To Flashcards QA (synth) Wiki To RCQA QA (synth) Nemotron Synth QA QA (synth) Math Meta Reasoning Thinking (synth) Code Meta Reasoning Thinking (synth) Program Verifiable Thinking (synth) OMR Rewrite FullThoughts Thinking (synth) QWQ Reasoning Traces Thinking (synth) General Reasoning Mix Thinking (synth) Gemini Reasoning Traces Thinking (synth) Llama Nemotron Reasoning Traces Thinking (synth) OpenThoughts2 Reasoning Traces Thinking (synth) Tulu 3 SFT Instruction (synth) Dolmino 1 Flan Instruction (synth) OLMOCR Science PDFs (High Q.) PDFs STEM Heavy Crawl Web pages Common Crawl (High Q.) Web pages Ingredients There were two ingredients used during stage 2 midtraining annealling of Olmo 3 32B. There were 2 versions of a 100B mix: Ingredient 1 100B tokens Mix composition: web pages, code, math/QA/thinking/instructio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy