WildJailbreak Dataset Card WildJailbreak is an open source synthetic safety training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent. 1. Vanilla Harmful : direct requests that could potentially elicit harmful responses from LMs. We apply GPT 4 to synthetically generate 50,050 vanilla harmful prompts across 13 risk categories, inspired by taxonomy from Weidinger et al. In addition, we pair the harmful prompts with helpful and detailed refusal responses, also synthetically generated with GPT 3.5. 2. Vanilla Benign : harmless prompts used to combat exaggerated safety, i.e., over refusal on benign queries. Motivated by the exaggerated safety categories in XSTest, we use GPT 4 to generate 50,050 prompts that superficially resemble unsafe prompts by keywords or discuss sensitive topics in non harmful ways. Similarly, we use GPT 3.5 to generate complying responses. 3. A…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy