π Overview: This is the official classifier for text behaviors in HarmBench. This model support standard (text) behaviors and contextual behaviors. π Example Notebook to use the classifier can be found here π» π¬ Chat Template: π‘Example usage: π Performances AdvBench GPTFuzz ChatGLM (Shen et al., 2023b) Llama Guard (Bhatt et al., 2023) GPT 4 (Chao et al., 2023) HarmBench (Ours) Standard 71.14 77.36 65.67 68.41 89.8 94.53 Contextual 67.5 71.5 62.5 64.0 85.5 90.5 Average (β) 69.93 75.42 64.29 66.94 88.37 93.19 Table 1: Agreement rates between previous metrics and classifiers compared to human judgments on our manually labeled validation set. Our classifier, trained on distilled data from GPT 4 0613, achieves performance comparable to GPT 4. π Citation:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy