OR Bench: An Over Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR Bench Hard 1K and Y axis shows the rejection rate on OR Bench Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line, with its slope determined by the quadratic regression coefficient of all the points, to represent the overall performance of all models. Overall Workflow Below is the overall workflow of our pipeline. We automate the process of producing seemingly toxic prompts that is able to produce updated prompts constantly. Detailed Model Performance Here are the radar plots of different model performances. The red area indicates the rejection rate of seemingly toxic prompts and the blue area indicates the acceptance rate of toxic prompts. In both cases, the plotted area is the smaller the better. Claude 2.1 Claude 2.1 Claude 3 Model Family Claude 3 Haiku Claude 3 Sonnet Claude 3 Opus Gemini Model Family Gemma 7b Gemini 1.0 pro G…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy