Model Card for OLMo 7B For transformers versions v4.40.0 or newer, we suggest using OLMo 7B HF instead. OLMo is a series of O pen L anguage Mo dels designed to enable the science of language models. The OLMo models are trained on the Dolma dataset. We release all code, checkpoints, logs (coming soon), and details involved in training these models. A new version of this model with a 24 point improvement on MMLU is available here . Model Details The core models released in this batch are the following: Size Training Tokens Layers Hidden Size Attention Heads Context Length OLMo 1B 3 Trillion 16 2048 16 2048 OLMo 7B 2.5 Trillion 32 4096 32 2048 OLMo 7B Twin 2T 2 Trillion 32 4096 32 2048 We are releasing many checkpoints for these models, for every 1000 traing steps. The naming convention is step1000 tokens4B . In particular, we focus on four revisions of the 7B models: Name HF Repo Model Revision Tokens Note OLMo 7B allenai/OLMo 7B main 2.5T The base OLMo 7B model OLMo 7B (not annealed) allenai/OLMo 7B step556000 tokens2460B 2.5T learning rate not annealed to 0 OLMo 7B 2T allenai/OLMo 7B step452000 tokens2000B 2T OLMo checkpoint at 2T tokens OLMo 7B Twin 2T allenai/OLMo 7B Twin 2T main…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy