OLMoE 1B 7B 0125 Instruct Release Documentation OLMoE 1B 7B 0125 Instruct January 2025 is post trained variant of the OLMoE 1B 7B January 2025 model, which has undergone supervised finetuning on an OLMo specific variant of the Tülu 3 dataset and further DPO training on this dataset, and finally RLVR training using this data. Tülu 3 is designed for state of the art performance on a diversity of tasks in addition to chat, such as MATH, GSM8K, and IFEval. Check out the OLMoE paper or Tülu 3 paper for more details! OLMo is a series of O pen L anguage Mo dels designed to enable the science of language models. These models are trained on the Dolma dataset. We are releasing all code, checkpoints, logs (coming soon), and associated training details. The core models released in this batch include the following: Stage OLMoE 1B 7B Base Model allenai/OLMoE 1B 7B 0125 SFT allenai/OLMoE 1B 7B 0125 SFT DPO allenai/OLMoE 1B 7B 0125 DPO Final Models (RLVR) allenai/OLMoE 1B 7B 0125 Instruct Reward Model (RM) allenai/OLMoE 1B 7B 0125 RM Model description Model type: A model trained on a mix of publicly available, synthetic and human created datasets. Language(s) (NLP): Primarily English License: Apac…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy