Qwen3.6 27B NVFP4 Mixed precision NVFP4 + FP8 quantization of Qwen/Qwen3.6 27B targeting native Blackwell (SM120) deployment — RTX 5090 32 GB, RTX 6000 Pro 96 GB. The original BF16 checkpoint needs ~52 GiB of VRAM. This build fits a single RTX 5090 32 GB at 32k context with usable KV cache, multi turn reasoning, and tool call support. Variants The repository hosts three branches, each tuned for a different deployment profile. Pick one with revision= in from pretrained or revision in the HF CLI. Branch lm head embed tokens Use case Status main BF16 BF16 Workflow / tool call / structured output Recommended default fp8 head FP8 BLOCK [128, 128] BF16 Free form text generation, more concurrency Stable fp8 head embed FP8 BLOCK [128, 128] FP8 BLOCK [128, 128] Maximum VRAM saving / max concurrency Lab only — see caveat below All three branches share the same inner layer quantization scheme (NVFP4 on the large Linears, FP8 per tensor on accuracy sensitive Linears, BF16 on normalization / GDN sub projections / MTP / visual tower). They differ only in the precision of the embedding boundary layers. VRAM and concurrency comparison Measured on a single RTX 5090 32 GB at max model len=32768 with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy