Dolma 3 Mix (6T) The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo 3 1125 32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl. For more information on Dolma, please see our original release here. Smaller Sample for Analysis Available! If you would like a smaller sample of this mix to examine and experiment with, which uses the same upsampling strategy, please see our 150B mix: https://huggingface.co/datasets/allenai/dolma3 mix 150B 1025. Licensing Information Dolma 3 mix is licensed under the Open Data Commons Attribution License v1.0 (ODC By). It is intended for research and educational use. For more information, please see our Responsible Use Guidelines. Citation Find the paper at: https://allenai.org/papers/olmo3
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy