Notice: Model quality improved significantly after adopting the MoQ strategy; however, because the MoQ strategy offers very few interchangeable Q4 K options, the NVFP4 model loses almost all its performance advantages under this approach. Consequently, unless new technology emerges to enhance NVFP4's quantization performance, we will likely remain on the v3 version for the foreseeable future. If model quality is your priority, I recommend my MoQ series, as these models offer superior quality: Jianqiao1/Qwen3.6 27B MTP MoQ GGUF Qwen3.6 27B NVFP4 MTP GGUF This is a GGUF quantization of Qwen3.6 27B MTP using a custom NVFP4 quantizer and a MoQ derived mixed tensor policy. The model was quantized with my customized llama.cpp build, but the output GGUF uses standard tensor types and is compatible with mainline llama.cpp builds that support NVFP4. This implementation incorporates ideas from michaelw9999's NVFP4 quantizer component in advanced gguf quantizer, uses the unsloth imatrix file for Qwen3.6 27B, and applies my rich+CJSO adaptive NVFP4 scale search with RSF lite. The tensors are stored using standard GGUF tensor types such as NVFP4 , IQ4 XS , Q5 K , Q8 0 , BF16 , and F32 . The mod…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy