TenyxChat: Language Model Alignment using Tenyx Fine tuning Introducing Llama 3 TenyxChat 70B, part of our TenyxChat series trained to function as useful assistants through preference tuning, using Tenyx's advanced fine tuning technology (VentureBeat article). Our model is trained using the Direct Preference Optimization (DPO) framework on the open source AI feedback dataset UltraFeedback. We fine tune Llama3 70B with our proprietary approach which shows an increase in MT Bench , without a drop in performance of the model on other benchmarks. Our approach aims to mitigate forgetting in LLMs in a computationally efficient manner, thereby enabling continual fine tuning capabilities without altering the pre trained output distribution. Llama 3 TenyxChat 70B was trained using eight A100s (80GB) for fifteen hours, with a training setup obtained from HuggingFaceH4 (GitHub). The MT Bench evaluation we perform follows the latest eval upgrade as PR'd here. This PR upgrades the evaluation from GPT 4 0613 to GPT 4 preview 0125 (latest version) as well as corrects and improves the quality of the reference answers for a subset of questions. These changes are required to correct the erroneous ra…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy