Read our How to Run Qwen3.6 NVFP4 Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. 1.79x faster throughput than other NVFP4 quants. This is an Unsloth NVFP4 quantized checkpoint calibrated on a mixture of our Unsloth dataset + UltraChat dataset. Works on a 32GB VRAM GPU. Benchmarks on 1xB200 128 concurrency. NVFP4 Accuracy Benchmarks For accuracy benchmarks, we conducted MMLU Pro, AIME 2025, GPQA for FP8, BF16, NVIDIA's NVFP4 and our NVFP4s we show our faster quants do similarly on all: Provider MMLU Pro GPQA AIME 2025 : : : Unsloth NVFP4 85.85 86.74 92.29 Unsloth NVFP4 Fast 85.58 87.75 91.67 NVIDIA NVFP4 85.60 87.12 91.88 FP8 85.75 86.74 93.12 BF16 85.75 86.36 92.50 Read all benchmarks in our NVFP4 blog vLLM Run Instructions To install vLLM in a separate venv: Then to serve the 35B Fast variant: Also do NOT use the Marlin backend since it's 2x slower use the native vLLM or cute DSL / CUTLASS / flashinfer trtllm backends! Model backend decode tok/s thr out tok/s nvidia 27B marlin (auto) 115.6 2,403 unsloth 27B cute DSL (auto) 125.9 6,863 nvidia 35B A3B marlin (auto) 240.8 8,721 unsloth 35B A3B cute DSL + trtllm (auto) 295.2 15,636 DGX Spark You must use the bel…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy