Recently re quantized with the latest version of advanced gguf quantizer tool. MTP included. Qwen3.6 27B RSF NVFP4 GGUF v4 This latest (6 June 2026) NVFP4 quant of Qwen3.6 27B keeps quality very high, while improving KLD, tail KLD metrics, top token stability, and probability error compared to the earlier releases. It is faster and smaller in size. It applies my RSF scale fitting technique to the Q K quants and uses a different tensor mix layout compared to previously. 6 June 2026: changed the MTP tensors to be NVFP4, ~130tk/s tg is now possible in some configurations on 5090. I am experimenting with a smaller verison of this model designed to be used on machines with 16GB of VRAM: Qwen3.6 27B NVFP4 SMALL MTP GGUF . The SMALL version will be slower than this model but should maintain similar quality, although I am still evaluating it. Feedback to improve this is appreciated. Results below on RTX 5090 with Wikitest: Model Size GiB pp512 tg128 pg32768,256 RSF NVFP4 15.27 5174.31 ± 1.30 76.61 ± 0.18 2751.93 ± 23.73 Quality Metric Earlier NVFP4 Improved RSF NVFP4 Change Mean PPL(Q) 7.128133 ± 0.047613 7.030348 ± 0.046636 1.37% lower Mean PPL(base) 6.9…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy