UltraFeedback Binarized using the Average of Preference Ratings (Cleaned) This dataset represents a new iteration on top of argilla/ultrafeedback binarized preferences , and is the recommended and preferred dataset by Argilla to use from now on when fine tuning on UltraFeedback . Read more about Argilla's approach towards UltraFeedback binarization at argilla/ultrafeedback binarized preferences/README.md . Differences with argilla/ultrafeedback binarized preferences Thanks to the recent issue identified by AllenAI related to the TruthfulQA contamination within the original UltraFeedback dataset due to some prompts being reused from the TruthfulQA dataset (used for benchmarking in the Open LLM Leaderboard from HuggingFace H4), we also decided to follow AllenAI's advice and remove those from the UltraFeedback dataset that we binarized using a completely different approach, which implied using the average of the preference ratings rather than the critique overall score, as HuggingFaceH4/ultrafeedback binarized did. Besides that, we also saw that not only the rows with the source=truthful qa were contamined (for obvious reasons), but also some coming from ShareGPT, so we also removed t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy