Read our How to Run Qwen3.6 NVFP4 Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. 1.56x faster throughput than other NVFP4 quants. This is an Unsloth NVFP4 quantized checkpoint calibrated on a mixture of our Unsloth dataset + UltraChat dataset. Works on a 32GB VRAM GPU. Benchmarks on 1xB200 128 concurrency. Use the 35B NVFP4 Fast version for 1.79x faster at a little less accuracy NVFP4 Accuracy Benchmarks For accuracy benchmarks, we conducted MMLU Pro, AIME 2025, GPQA for FP8, BF16, NVIDIA's NVFP4 and our NVFP4s we show our faster quants do similarly on all: Provider MMLU Pro GPQA AIME 2025 : : : Unsloth NVFP4 85.85 86.74 92.29 Unsloth NVFP4 Fast 85.58 87.75 91.67 NVIDIA NVFP4 85.60 87.12 91.88 FP8 85.75 86.74 93.12 BF16 85.75 86.36 92.50 Read all benchmarks in our NVFP4 blog vLLM Run Instructions To install vLLM in a separate venv: Then to serve the 35B variant: DGX Spark You must use the below or you will get 2x slower inference! Also do NOT use the Marlin backend since it's 2x slower use the native vLLM or cute DSL / CUTLASS / flashinfer trtllm backends! Model backend decode tok/s thr out tok/s nvidia 27B marlin (auto) 115.6 2,403 unsloth 27B cute DSL (auto…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy