Qwen3.6 27B NVFP4 Model Card Overview This is an NVFP4 quantized version of Qwen/Qwen3.6 27B by Lna Lab, using custom Blackwell NVFP4 GEMM kernels. This is the first NVFP4 release in our Qwen3.6 27B family — compressed tensors format, vision tower preserved, no MTP head. ~35K downloads since release. For new deployments we strongly recommend the faster siblings below unless you have a reason to stay on compressed tensors . Key Compression Stats: Original Size: 55.6 GB Quantized Size: 19.7 GB (0.35x compression) Vision Tower: Preserved in BF16 Hardware: Runs on a single NVIDIA Blackwell GPU Faster siblings — modelopt + MTP format Verified throughput at the same production launch (single 1× RTX PRO 6000 Blackwell, vLLM 0.19.1rc1, 256K context, KV FP8, max num seqs 2): Repo Format MTP Single tok/s 2 parallel agg tok/s vs this repo Qwen3.6 27B NVFP4 (this) compressed tensors ❌ 58 (M / L) 119 (M / L) 1.0× (baseline) Qwen3.6 27B Text NVFP4 MTP modelopt ✅ n=3 98 / 100 189 / 207 1.67× / 1.74× Carnice V2 27b NVFP4 TEXT MTP modelopt ✅ n=3 98 / 102 193 / 194 1.68× / 1.63× Huihui Qwen3.6 27B abliterated NVFP4 TEXT MTP modelopt ✅ n=3 96 / 101 203 / 183 1.65× / 1.54× Huihui Qwen3.6 27B abliterat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy