Qwopus3.6 27B Coder NVFP4 MTP GGUF NVFP4 (Blackwell native FP4) quantized GGUF of Jackrong/Qwopus3.6 27B Coder MTP for llama.cpp with Multi Token Prediction (MTP) speculative decoding support. Quantization Details Attribute Value Source Q8 0 GGUF (28 GB, 8.50 BPW) Output Mixed precision NVFP4 (15 GB, 4.60 BPW) Size reduction 46% (28 GB → 15 GB) NVFP4 tensors 311 (attn q, attn k, attn v, attn qkv, attn output, ffn down, ffn gate, ffn up) Q4 K tensors 194 (attn gate, ssm alpha, ssm beta, ssm out, nextn.eh proj, token embd) Q4 K S tensors 1 (output.weight) F32 tensors 360 (norms, biases, SSM state) Total tensors 866 Tensor Mapping Strategy Following Unsloth's NVFP4 approach for Qwen3.6 27B hybrid Mamba2 Transformer models: NVFP4 → 8 large weight tensor patterns that dominate model size and bandwidth (attention projections + FFN weights) Q4 K → Smaller weights (SSM parameters, MTP head projection, token embeddings) — preserves quality where tensor dimensions are small F32 → Norms, biases, and SSM state — must remain full precision for numerical stability This mapping matches the reference Qwen3.6 27B NVFP4 MTP quantization exactly. Performance NVIDIA DGX Spark (GB10, ARM64, 128 GB unif…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy