Suzume ORPO [[Paper]](https://arxiv.org/abs/2405.18952) [[Dataset]](https://huggingface.co/datasets/lightblue/mitsu) This is Suzume ORPO, an ORPO trained fine tune of the lightblue/suzume llama 3 8B multilingual model using our lightblue/mitsu dataset. We have trained several versions of this model using ORPO and so recommend that you use the best performing model from our tests, lightblue/suzume llama 3 8B multilingual orpo borda half. Note that this model has a non commerical license as we used the Command R and Command R+ models to generate our training data for this model (lightblue/mitsu). We are currently working on a developing a commerically usable model, so stay tuned for that! Model list We have ORPO trained the following models using different proportions of the lightblue/mitsu dataset: Trained on the top/bottom responses of all prompts in the dataset: lightblue/suzume llama 3 8B multilingual orpo borda full Trained on the top/bottom responses of the prompts of the 75\% most consistently ranked responses in the dataset: lightblue/suzume llama 3 8B multilingual orpo borda top75 Trained on the top/bottom responses of the prompts of the 50\% most consistently ranked respons…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy