NVFP4 Quantized RedHatAI/Qwen3.6 35B A3B NVFP4 This is a preliminary version (and subject to change) of NVFP4 quantized Qwen/Qwen3.6 35B A3B model. The model has both weights and activations quantized to NVFP4 format with vllm project/llm compressor. It is compatible and tested against vllm main. Deploy it with: vllm serve RedHatAI/Qwen3.6 35B A3B NVFP4 reasoning parser qwen3 moe backend flashinfer cutlass . If you have hardware with more compute than memory bandwidth, you may prefer this MoE variant for performance reasons. Creation Script: Run this script with LLM Compressor main and latest transformers. Evaluation This model was evaluated on GSM8K Platinum, MMLU Pro, IFEval, Math 500, GPQA Diamond, AIME 25, and LiveCodeBench v6 using lm evaluation harness and lighteval, served with vLLM using language model only . Accuracy Benchmark Qwen/Qwen3.6 35B A3B RedHatAI/Qwen3.6 35B A3B NVFP4 Recovery (%) GSM8k Platinum (0 shot) 95.73 96.08 100.37 IfEval (0 shot) 93.09 92.45 99.31 AIME 2025 92.92 91.25 98.21 GPQA diamond 84.51 84.68 100.20 Math 500 84.80 85.00 100.24 Lcb Codegeneration V6 77.33 74.67 96.55 MMLU Pro Chat 85.32 84.70 99.28 BFCLv4 Overall 57.83 56.10 97.01 BFCLv4 Single Tur…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy