distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca dpo pairs. The original dataset has been used by 100s of open source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open source LLMs and the community, we spent a few hours improving it with distilabel. This was our main intuition: the original dataset just assumes gpt4/3.5 turbo are always the best response. We know from UltraFeedback that's not always the case. Moreover, DPO fine tuning benefits from the diversity of preference pairs. Additionally, we have added a new column indicating whether the question in the dataset is part of the train set of gsm8k (there were no examples from the test set). See the reproduction section for more details. Using this dataset This dataset is useful for preference tuning and we recommend using it instead of the original. It's already prepared in the "standard" chosen, rejected format with additional information for further filtering and experimentation. The main changes are: 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy