Qwen3.6 27B NVIDIA NVFP4 MTP GGUF GGUF conversion of nvidia/Qwen3.6 27B NVFP4, preserving NVIDIA NVFP4 tensors, with MTP speculative decoding and a BF16 vision projector. Benchmarked on an RTX 5090 with llama benchy . Highlights NVFP4 preserved : 193 NVFP4 tensors are preserved from NVIDIA's ModelOpt quantized checkpoint. Q4 attention : attention ( q/k/v/o ) and the linear attention / DeltaNet projections are quantized to Q4 K (down from the original FP8) to keep this build compact. For those layers kept at Q8 0 for better accuracy — about +3 GB — see the Q8attn variant. MTP included : the GGUF keeps the extra MTP layer for draft mtp speculative decoding. Vision supported : includes a BF16 mmproj file for image input. RTX 5090 tested : measured with llama benchy using MTP depth d=3 . Provenance Component Source Base model Qwen/Qwen3.6 27B NVFP4 source checkpoint nvidia/Qwen3.6 27B NVFP4 Runtime target llama.cpp Files File Size Description : Qwen3.6 27B NVIDIA NVFP4 MTP.gguf 14.66 GiB / 15,747,650,944 bytes Main GGUF. NVFP4 tensors preserved; MTP layer included. mmproj Qwen3.6 27B NVIDIA NVFP4 BF16.gguf ~888 MiB / 931,146,304 bytes BF16 vision projector for image input. llama.cpp ex…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy