Code Leaderboard Prior Preference Sets Results Paper Reward Bench Evaluation Dataset Card The RewardBench evaluation dataset evaluates capabilities of reward models over the following categories: 1. Chat : Includes the easy chat subsets (alpacaeval easy, alpacaeval length, alpacaeval hard, mt bench easy, mt bench medium) 2. Chat Hard : Includes the hard chat subsets (mt bench hard, llmbar natural, llmbar adver neighbor, llmbar adver GPTInst, llmbar adver GPTOut, llmbar adver manual) 3. Safety : Includes the safety subsets (refusals dangerous, refusals offensive, xstest should refuse, xstest should respond, do not answer) 4. Reasoning : Includes the code and math subsets (math prm, hep cpp, hep go, hep java, hep js, hep python, hep rust) The RewardBench leaderboard averages over these subsets and a final category from prior preference data test sets including Anthropic Helpful, Anthropic HHH in BIG Bench, Stanford Human Preferences (SHP), and OpenAI's Learning to Summarize data. The scoring for RewardBench compares the score of a prompt chosen pair to a prompt rejected pair. Success is when the chosen score is higher than rejected. In order to create a representative, single evaluat…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy