vllm (pretrained=/root/autodl tmp/Qwen3 32B abliterated,add bos token=true,max model len=3096,dtype=bfloat16,trust remote code=true,tensor parallel size=4,gpu memory utilization=0.8), gen kwargs: (None), limit: 250.0, num fewshot: 5, batch size: auto Tasks Version Filter n shot Metric Value Stderr : : : : gsm8k 3 flexible extract 5 exact match ↑ 0.900 ± 0.0190 strict match 5 exact match ↑ 0.896 ± 0.0193 vllm (pretrained=/root/autodl tmp/Qwen3 32B abliterated,add bos token=true,max model len=3096,dtype=bfloat16,trust remote code=true,tensor parallel size=4,gpu memory utilization=0.8), gen kwargs: (None), limit: 500.0, num fewshot: 5, batch size: auto Tasks Version Filter n shot Metric Value Stderr : : : : gsm8k 3 flexible extract 5 exact match ↑ 0.852 ± 0.0159 strict match 5 exact match ↑ 0.840 ± 0.0164 Groups Version Filter n shot Metric Value Stderr : : : mmlu 2 none acc ↑ 0.7988 ± 0.0131 humanities 2 none acc ↑ 0.7897 ± 0.0269 other 2 none acc ↑ 0.7590 ± 0.0298 social sciences 2 none acc ↑ 0.8722 ± 0.0252 stem 2 none acc ↑ 0.7860 ± 0.0230 vllm (pretrained=/root/autodl tmp/Qwen3 32B abliterated awq,add bos token=true,max model len=3096,dtype=bfloat16,trust remote code=true,tensor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy