This model is a Slerp Merge of cookinai/CatMacaroni Slerp and mncai/mistral 7b dpo v5. Evaluation Results HuggingFace Leaderboard Average ARC HellaSwag MMLU TruthfulQA Winogrande GSM8K 73.1 69.62 87.09 64.81 62.82 81.45 72.78 The model did achieve an improvement in TruthfulQA over cookinai/CatMacaroni Slerp and GSM8K over mncai/mistral 7b dpo v5 which was the goal of the merge leading to an average score that was a better than both. It is unclear why the TruthfulQA metric is still meaningfully lower than the base mncai/mistral 7b dpo v5 . Training Details .yaml file for mergekit Bias, Risks, and Limitations The model has not been evaluated for safety and is only intended for research and experiments.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy