Qwen3.6 27B NVFP4 NVFP4 (W4A4) quantization of Qwen/Qwen3.6 27B, produced with NVIDIA Model Optimizer. The vision encoder is preserved in BF16 so multimodal (image + video) capability is intact; only the 27B language backbone is quantized. Weights size: 19 GB (vs ~54 GB BF16, ~2.8× compression) Hardware target: NVIDIA Blackwell (RTX 5090, B100/B200) — native FP4 tensor cores Calibration: mixed image text + text only , 128 + 128 samples (see below) KV cache: FP8 (E4M3, fp8 cast — amax set to FP8 range without data driven calibration of K/V tensors). Matches the convention used by nvidia/Qwen3 32B NVFP4 , nvidia/Qwen3.5 397B A17B NVFP4 , nvidia/Gemma 4 31B IT NVFP4 . What's quantized Component Format Language model linear weights/activations (full + linear attention transformer blocks) NVFP4 (W4A4, group size 16, FP8 E4M3 scales) KV cache FP8 (E4M3) Vision encoder ( model.visual. , 27 SigLIP style blocks) BF16 (untouched) Vision to LM projector + image/video embeddings BF16 lm head BF16 Linear attention conv1d (48 of 64 layers in this hybrid model) BF16 MTP head ( mtp , mtp.layers.0 ) BF16 Routers / mlp.gate. BF16 The exclusion list is recorded in hf quant config.json ( exclude modul…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy