Self Play Preference Optimization for Language Model Alignment (https://arxiv.org/abs/2405.00675) Mistral7B PairRM SPPO Iter3 This model was developed using Self Play Preference Optimization at iteration 3, based on the mistralai/Mistral 7B Instruct v0.2 architecture as starting point. We utilized the prompt sets from the openbmb/UltraFeedback dataset, splited to 3 parts for 3 iterations by snorkelai/Snorkel Mistral PairRM DPO Dataset. All responses used are synthetic. This is the model reported in the paper , with K=5 (generate 5 responses per iteration). We attached the Arena Hard eval results in this model page. Links to Other Models Mistral7B PairRM SPPO Iter1 Mistral7B PairRM SPPO Iter2 Mistral7B PairRM SPPO Iter3 Mistral7B PairRM SPPO Model Description Model type: A 7B parameter GPT like model fine tuned on synthetic datasets. Language(s) (NLP): Primarily English License: Apache 2.0 Finetuned from model: mistralai/Mistral 7B Instruct v0.2 AlpacaEval Leaderboard Evaluation Results Model LC. Win Rate Win Rate Avg. Length : : : : : : Mistral7B PairRM SPPO Iter 1 24.79 23.51 1855 Mistral7B PairRM SPPO Iter 2 26.89 27.62 2019 Mistral7B PairRM SPPO Iter 3 28.53 31.02 2163 Mistral7B…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy