Introduction GitHub Repo UltraRM 13b UltraCM 13b UltraFeedback is a large scale, fine grained, diverse preference dataset , used for training powerful reward models and critic models. We collect about 64k prompts from diverse resources (including UltraChat, ShareGPT, Evol Instruct, TruthfulQA, FalseQA, and FLAN). We then use these prompts to query multiple LLMs (see Table for model lists) and generate 4 different responses for each prompt, resulting in a total of 256k samples. To collect high quality preference and textual feedback, we design a fine grained annotation instruction, which contains 4 different aspects, namely instruction following , truthfulness , honesty and helpfulness . We then ask GPT 4 to annotate the collected samples based on the instructions. Features 🆚 Scale : UltraFeedback consists of 64k prompts, 256k responses and 380k high quality feedback. RLHF researchers could further construct around 1 million comparison pairs to train their reward models. 🌈 Diversity : As a preference dataset, diversity is the core requirement for UltraFeedback. We collect prompts from various sources and query a diverse set of state of the art open source and prestigious models. T…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy