Model Card for OLMo 7B OLMo is a series of O pen L anguage Mo dels designed to enable the science of language models. The OLMo models are trained on the Dolma dataset. We release all code, checkpoints, logs (coming soon), and details involved in training these models. This model has been converted from allenai/OLMo 7B for the Hugging Face Transformers format. Model Details The core models released in this batch are the following: Size Training Tokens Layers Hidden Size Attention Heads Context Length OLMo 1B 3 Trillion 16 2048 16 2048 OLMo 7B 2.5 Trillion 32 4096 32 2048 OLMo 7B Twin 2T 2 Trillion 32 4096 32 2048 We are releasing many checkpoints for these models, for every 1000 training steps. These have not yet been converted into Hugging Face Transformers format, but are available in allenai/OLMo 7B. Model Description Developed by: Allen Institute for AI (AI2) Supported by: Databricks, Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University, AMD, CSC (Lumi Supercomputer), UW Model type: a Transformer style autoregressive language model. Language(s) (NLP): English License: The code and model are released under Apache 2.0. Contact: Technical inq…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy