⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3 mix 6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use our 32B training mix, which utilizes the same sampling strategy and is complete with olmOCR science pdfs. Dolma 3 Mix (6T) The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo 3 1025 7B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl. For more information on Dolma, please see our original release here. Dataset Sources Source Sizes This dataset contains the full mix of documents used to train Olmo 3 7B. Source Doc Type Tokens Bytes (uncompressed) Documents License common crawl web pages 4.51T 18.0TB 3.15B ODC BY olmocr science pdfs ac…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy