Self Play Preference Optimization for Language Model Alignment (https://arxiv.org/abs/2405.00675) Llama 3 Instruct 8B SPPO Iter2 This model was developed using Self Play Preference Optimization at iteration 2, based on the meta llama/Meta Llama 3 8B Instruct architecture as starting point. We utilized the prompt sets from the openbmb/UltraFeedback dataset, splited to 3 parts for 3 iterations by snorkelai/Snorkel Mistral PairRM DPO Dataset. All responses used are synthetic. Links to Other Models Llama 3 Instruct 8B SPPO Iter1 Llama 3 Instruct 8B SPPO Iter2 Llama 3 Instruct 8B SPPO Iter3 Model Description Model type: A 8B parameter GPT like model fine tuned on synthetic datasets. Language(s) (NLP): Primarily English License: Apache 2.0 Finetuned from model: meta llama/Meta Llama 3 8B Instruct AlpacaEval Leaderboard Evaluation Results Model LC. Win Rate Win Rate Avg. Length : : : : : : Llama 3 8B SPPO Iter1 31.73 31.74 1962 Llama 3 8B SPPO Iter2 35.15 35.98 2021 Llama 3 8B SPPO Iter3 38.77 39.85 2066 Open LLM Leaderboard Evaluation Results Results are reported by using lm evaluation harness v0.4.1 arc challenge truthfulqa mc2 winogrande gsm8k hellaswag mmlu average Llama 3 8B SPPO Ite…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy