RTX 5090 Windows native GGUF comparison Qwen3.6 27B NVFP4 MTP GGUF Native llama.cpp conversion of unsloth/Qwen3.6 27B NVFP4. This is a separate artifact from the existing NVIDIA source file in this repository. Files File Source Format qwen3.6 27b nvfp4.gguf NVIDIA NVFP4 Canonical NVFP4 download; bundled MTP preserved qwen3.6 27b nvfp4 unsloth.gguf Unsloth NVFP4 Native NVFP4 FFNs, source FP8 tensors stored as Q8 0, bundled MTP preserved Size: 21.58 GiB. SHA256: DDCE600959ED99C16092EF24E8885AE18FD0D0B8A53EF92976BE1A395194FBD3 . The source checkpoint uses a mixed compressed tensors layout. The conversion keeps packed NVFP4 FFN weights native and stores the source FP8 weights as Q8 0 to leave usable KV cache headroom on an RTX 5090. The source mtp. block is preserved for draft mtp speculative decoding. RTX 5090 Result Windows 11, RTX 5090, llama.cpp b9851, ctx=200000 , q4 0 target/draft KV, draft mtp n=2 , no thinking, identical BookContext prompt fixture, and one measured 1024 token completion. Source 10k decode tok/s 200k decode tok/s VRAM after Temperature after : : : : NVIDIA source GGUF 69.0 42.2 30.8 GiB 52 C / 61 C Unsloth source GGUF 72.8 44.1 26.1 GiB 46 C / 60 C decode tok/s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy