This Qwen3.5 9B model was quantized with NVFP4 with MTP support, using my soon to be released NVFP4 GGUF quantizer. This autotunes the model to reduce ppl and kld as much as possible, and selects an optimal llama.cpp tensor distribution. This keeps the model size down to just 5.66GB (previously 6.21GB) while still improving quality, and MTP increases speed. Updated results with better ppl/kld distribution: Previous results of this model (older quantizer): Compare the original Qwen3.5 9B NVFP4(made via ModelOpt):
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy