Dataset Card for PKU SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU Alignment Team or any of its members. [🏠 Homepage] [🤗 Single Dimension Preference Dataset] [🤗 Q A Dataset] [🤗 Prompt Dataset] Citation If PKU SafeRLHF has contributed to your work, please consider citing our research: If you encounter any issues with our dataset, please contact us through the HuggingFace Discussion. Dataset Summary This dataset is a sibling project of PKU SafeRLHF v0 and BeaverTails. We provide a high quality dataset consisting of 83.4K preference entries, which is annotated across two dimensions: harmlessness and helpfulness. Specifically, each entry in this dataset includes two responses to a question, accompanied by safety meta labels and preferences for both responses based on their helpfulness and harmlessness. For a more fine grained labeling of Q A pairs in this dataset, see PKU SafeRLHF QA. In this work, we performed SFT on Llama2 7B and Llama3 8B with Alpaca 52K dataset, resulting in Alpaca2 7…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy