⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool , pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3 mix 6T 1025 . Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general, please see our original release here. A Note on the Dolma 3 Pool: Source Links The dolma 3 pool contains documents for Common Crawl (web) and olmOCR Science PDFs only . To access the documents from the remaining sources in this pool, follow the source links below: Common Crawl : Current repository olmOCR Science PDFs : Current repository StackEdu : https://huggingface.co/datasets/HuggingFaceTB/stack edu arXiv : https://huggingface.co/datasets/togethercomputer/RedPajama Data 1T FineMath 3+ : https://huggingface.co/datasets/HuggingFaceTB/finemath Wikipedia & Wikibooks : https://huggingface.co/datasets/allenai/dolma (dolma v1.7) Dataset Sources This dataset contains the full pool of documents considered to train th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy