Nvidia Qwen3.6 27B NVFP4 GGUF Quantized GGUF versions of nvidia/Qwen3.6 27B NVFP4. These were generated using llama.cpp (b9859). Nvidia Qwen3.6 27B NVFP4 A.gguf : The FFN and attention layers are NVFP4 quantized. The MTP draft layer is preserved at BF16. This was generated from BF16 Attn using a modified llama quantize to re quantize the BF16 upcast attention layers to NVFP4. Nvidia Qwen3.6 27B NVFP4 Q8.gguf : The NVFP4 FFN layers are preserved, while the FP8 attention layers are converted to Q8. This was generated using convert hf to gguf.py . Nvidia Qwen3.6 27B NVFP4 BF16 Attn.gguf : The NVFP4 FFN layers are preserved, while the FP8 attention layers are upcast to BF16. This is the default conversion for BF16 because GGUF files do not support FP8. This was generated using convert hf to gguf.py . Quantizations provided File Quantization Size Nvidia Qwen3.6 27B NVFP4 A.gguf NVFP4 FFN and attention 17.9 GB Nvidia Qwen3.6 27B NVFP4 Q8.gguf NVFP4 FFN, Q8 attention 21.5 GB Nvidia Qwen3.6 27B NVFP4 BF16 Attn.gguf NVFP4 FFN, BF16 attention 28.2 GB Perplexity test I tested perplexity using llama perplexity and Salesforce's wikitext 2 raw v1. File Ctx PPL Nvidia Qwen3.6 27B NVFP4 A.gguf 512…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy