Model Card for OLMo 7B July 2024 OLMo is a series of O pen L anguage Mo dels designed to enable the science of language models. The OLMo models are trained on the Dolma dataset. We release all code, checkpoints, logs, and details involved in training these models. Model Details The core models released in this batch are the following: Size Training Tokens Layers Hidden Size Attention Heads Context Length OLMo 1B July 2024 3.05 Trillion 16 2048 16 4096 OLMo 7B July 2024 2.75 Trillion 32 4096 32 4096 [Coming soon] We are releasing many checkpoints for these models, for every 1000 training steps. The naming convention is stepXXX tokensYYYB . These checkpoints are already available at OLMo 7B April 2024 and will be copied here soon. To load a specific model revision with HuggingFace, simply add the argument revision : All revisions/branches are listed in the file revisions.txt . Or, you can access all the revisions for the models via the following code snippet: Model Description Developed by: Allen Institute for AI (AI2) Supported by: Databricks, Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University, AMD, CSC (Lumi Supercomputer), UW Model type: a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy