Llama 3.1 Tulu 3 8B SFT Tülu3 is a leading instruction following model family, offering fully open source data, code, and recipes designed to serve as a comprehensive guide for modern post training techniques. Tülu3 is designed for state of the art performance on a diversity of tasks in addition to chat, such as MATH, GSM8K, and IFEval. Model description Model type: A model trained on a mix of publicly available, synthetic and human created datasets. Language(s) (NLP): Primarily English License: Llama 3.1 Community License Agreement Finetuned from model: meta llama/Llama 3.1 8B Model Sources Training Repository: https://github.com/allenai/open instruct Eval Repository: https://github.com/allenai/olmes Paper: https://arxiv.org/abs/2411.15124 Demo: https://playground.allenai.org/ Model Family Stage Llama 3.1 8B Llama 3.1 70B Base Model meta llama/Llama 3.1 8B meta llama/Llama 3.1 70B SFT allenai/Llama 3.1 Tulu 3 8B SFT allenai/Llama 3.1 Tulu 3 70B SFT DPO allenai/Llama 3.1 Tulu 3 8B DPO allenai/Llama 3.1 Tulu 3 70B DPO Final Models (RLVR) allenai/Llama 3.1 Tulu 3 8B allenai/Llama 3.1 Tulu 3 70B Reward Model (RM) allenai/Llama 3.1 Tulu 3 8B RM (Same as 8B) Stage Llama 3.1 405B Base Mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy