Model Summary This model is an intermediate training checkpoint during post training, after the Supervised Fine Tuning (SFT) step. For best performance, we recommend you use the OLMoE Instruct version. Paper : https://arxiv.org/abs/2409.02060 Pretraining Checkpoints, Code, Data and Logs. SFT (Supervised Fine Tuning) Checkpoints, Code, Data and Logs. DPO/KTO (Direct Preference Optimization/Kahneman Tversky Optimization) , Checkpoints, Preference Data, DPO code, KTO code and Logs. Branches: main : Instruction tuned / supervised finetuned (SFT) model of https://hf.co/allenai/OLMoE 1B 7B 0924 ( main branch) load balancing : Ablation with load balancing loss during SFT non annealed : Ablation starting from the checkpoint prior to annealing (branch step1200000 tokens5033B of https://hf.co/allenai/OLMoE 1B 7B 0924) rather than the annealed checkpoint (branch main of https://hf.co/allenai/OLMoE 1B 7B 0924) Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy