PandaBench PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies. The PandaGuard framework architecture illustrating the end to end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges. Dataset Description This repository contains the benchmark results from extensive evaluations of various LLMs against different jailbreak attacks and defense mechanisms. The dataset enables researchers to: 1. Compare the effectiveness of different defense mechanisms against various attack methods 2. Analyze the safety capability tradeoffs of defensive systems 3. Evaluate the robustness of different LLMs to jailbreak attempts 4. Develop and test new defense algorithms with consistent evaluation metrics PandaBench builds comprehensive benchmarks for LLM/attack/defense/evaluation (a) Attack Success Rate vs. release date for various LLMs. (b) ASR across different harm categories with and without defense mechanisms. (c) Overall ASR for all evaluated LLMs with and without defense mechanisms. Dataset Structure The benchmark dataset is…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy