Evaluations are produced with https://github.com/neuralmagic/GuardBench and vLLM as an inference engine. Evaluations are obtained with vllm==0.15.0 and bug fixes from this PR. Dataset meta llama/Llama Guard 4 12B F1 RedHatAI/Llama Guard 4 12B quantized.w4a16 (this model) F1 F1 Recovery % meta llama/Llama Guard 4 12B Recall RedHatAI/Llama Guard 4 12B quantized.w4a16 Recall Recall Recovery % : : : : : : : : : : : : AART 0.874 0.865 98.97 0.776 0.761 98.07 AdvBench Behaviors 0.964 0.968 100.41 0.931 0.938 100.75 AdvBench Strings 0.83 0.823 99.16 0.709 0.699 98.59 BeaverTails 330k 0.732 0.727 99.32 0.591 0.584 98.82 Bot Adversarial Dialogue 0.513 0.499 97.27 0.376 0.361 96.01 CatQA 0.932 0.927 99.46 0.873 0.864 98.97 ConvAbuse 0.241 0.248 102.9 0.148 0.156 105.41 DecodingTrust Stereotypes 0.591 0.54 91.37 0.419 0.37 88.31 DICES 350 0.118 0.118 100 0.063 0.063 100 DICES 990 0.219 0.226 103.2 0.135 0.135 100 Do Anything Now Questions 0.746 0.74 99.2 0.595 0.587 98.66 DoNotAnswer 0.546 0.539 98.72 0.376 0.368 97.87 DynaHate 0.603 0.587 97.35 0.481 0.459 95.43 HarmEval 0.56 0.571 101.96 0.389 0.4 102.83 HarmBench Behaviors 0.959 0.954 99.48 0.922 0.912 98.92 HarmfulQ 0.86 0.857 99.65 0.755…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy