Model Card for OLMo 1B July 2024 OLMo 1B July 2024 is the latest version of the original OLMo 1B model rocking a 4.4 point increase in HellaSwag, among other evaluations improvements, from an improved version of the Dolma dataset and staged training. This version is for direct use with HuggingFace Transformers from v4.40 on. OLMo is a series of O pen L anguage Mo dels designed to enable the science of language models. The OLMo models are trained on the Dolma dataset. We release all code, checkpoints, logs, and details involved in training these models. Model Details The core models released in this batch are the following: Size Training Tokens Layers Hidden Size Attention Heads Context Length OLMo 1B July 2024 3.05 Trillion 16 2048 16 4096 OLMo 7B July 2024 2.75 Trillion 32 4096 32 4096 [Coming soon] We are releasing many checkpoints for these models, for every 1000 training steps. The naming convention is stepXXX tokensYYYB . To load a specific model revision with HuggingFace, simply add the argument revision : All revisions/branches are listed in the file revisions.txt . Or, you can access all the revisions for the models via the following code snippet: Model Description Develope…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy