Dataset Card for UltraFeedback Binarized Dataset Description This is a pre processed version of the UltraFeedback dataset and was used to train Zephyr 7Β β, a state of the art chat model at the 7B parameter scale. The original UltraFeedback dataset consists of 64k prompts, where each prompt is accompanied with four model completions from a wide variety of open and proprietary models. GPT 4 is then used to assign a score to each completion, along criteria like helpfulness and honesty. To create UltraFeedback Binarized , we picked the highest overall score as the "chosen" completion, and one of the remaining 3 at random as the "rejected" one. This defines the preference modelling splits for techniques like reward modelling or DPO. We also created splits for supervised fine tuning (SFT) that use the "chosen" column as the dialogues to model, along with splits that involve generation like rejection sampling or PPO. For details on the dataset processing, see the accompanying script. Dataset Structure Usage To load the dataset, run: Note: after the release of Zephyr 7b β, the team at Argilla noted that there were a few hundred completions with the incorrect label. Similarly, members of t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy