Qwen3.6 27B VLM NVFP4 MTP Internal testing artifact. Used for development. Not evaluated for production use. No quality benchmarks beyond a single GSM8K 50 sanity gate (47/50). Passes vibe check. NVFP4 (modelopt format) quantization of Qwen/Qwen3.6 27B with the MTP draft head retained in BF16, vision tower retained in BF16, lm head retained in BF16. Contents model.safetensors (~20.6 GB): single shard NVFP4 packed body weights ( uint8 packed + per block float8 e4m3fn weight scale + per tensor float32 weight scale 2 ) BF16 vision tower (333 model.visual. tensors) BF16 MTP head (15 mtp. tensors) BF16 lm head.weight BF16 linear attn.conv1d and in proj projections config.json — quantization config.ignore lists the 65 entries kept in BF16 (50 vision blocks, 15 MTP modules) hf quant config.json — modelopt metadata chat template.jinja — froggeric/Qwen Fixed Chat Templates (see Patches) tokenizer.json , tokenizer config.json , preprocessor config.json , video preprocessor config.json , generation config.json Input checkpoint size: 55.6 GB BF16 → output 20.6 GB (0.37×). Base recipe The 5 step graft procedure is from lna lab/GGUF to NVFP4 SM120 — credit to Tonoken / LNA LAB. Recipe doc: docs/…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy