Model Details Model Card for Olmo 3 7B We introduce Olmo 3, a new family of 7B and 32B models. This suite includes Base, Instruct, and Think variants. The Base models were trained using a staged training approach. Olmo is a series of O pen l anguage mo dels designed to enable the science of language models. These models are trained on the Dolma 3 dataset. We are releasing all code, checkpoints, and associated training details. Size Training Tokens Layers Hidden Size Q Heads KV Heads Context Length OLMo 3 7B 5.93 Trillion 32 4096 32 32 65,536 OLMo 3 32B 5.50 Trillion 64 5120 40 8 65,536 The core models released in this batch include the following: Stage Olmo 3 7B Think Olmo 3 32B Think Olmo 3 7B Instruct Base Model Olmo 3 7B Olmo 3 32B Olmo 3 7B SFT Olmo 3 7B Think SFT Olmo 3 32B Think SFT Olmo 3 7B Instruct SFT DPO Olmo 3 7B Think DPO Olmo 3 32B Think DPO Olmo 3 7B Instruct DPO Final Models (RLVR) Olmo 3 7B Think Olmo 3 32B Think Olmo 3 7B Instruct Installation Olmo 3 is supported in transformers v4.57.0 or higher: Inference You can use OLMo with the standard HuggingFace transformers library: For faster performance, you can quantize the model using the following method: The quantiz…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy