Qwen3.6 27B NVFP4 NVFP4 quantized version of Qwen/Qwen3.6 27B . 55.6 GB → 20.6 GB (0.37x) with vision tower and MTP draft head preserved in BF16. Tested on NVIDIA DGX Spark (GB10, SM 121). NVFP4 Quantization Details Base model Qwen/Qwen3.6 27B Quantization NVFP4 (W4A4 — weights FP4, activations FP4, scales FP8) Format compressed tensors (native vLLM support) Tool vllm project/llm compressor Calibration nvidia/Nemotron Post Training Dataset v2 (512 samples) Container eugr/spark vllm docker Size 20.6 GB (quantized shard + BF16 MTP shard) Requires NVIDIA Blackwell GPU (SM 120+), vLLM = 0.19 Recipe What's Quantized / What's Not Quantized (NVFP4): All Linear layers in the language model Kept in BF16: lm head , all vision layers ( model.visual. ), MLP gates, MTP draft head ( mtp. ) MTP Speculative Decoding The MTP draft head ( mtp.fc , mtp.layers.0. ) is kept in BF16 and shipped as a separate model mtp bf16.safetensors shard. Quantizing the draft head to FP8/FP4 lowers acceptance rate; BF16 is the typical choice for Qwen3.5 series NVFP4 checkpoints. Enable via vLLM: Evaluation humaneval instruct chat (lm evaluation harness, 0 shot, extract code filter): Model Metric Value Stderr : : Qwen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy