This is the second release of the NVFP4 version of Jackrong's Qwopus3.6 27B v2 MTP GGUF. 9 June 2026: Fixed MTP heads to NVFP4; performance with MTP is much improved (108tk/s tg) Please note, I am not affiliated; this is my own quantization effort made with my experimental work in progress advanced gguf quantizer. More evaluatlions will be underway, this page will be updated when those are complete. Feedback on how to improve the quantizer/this quantization is appreciated. For improved performance and quality, try my llama.cpp NVFP4 Repack with MXFP6 from: https://github.com/michaelw9999/llama.cpp/tree/nvfp4repack mxfp6 cuda These branches are updated regularly. NVFP4 repack preloads all tensors into a CUDA tile to boost speed. It is a tiny bit slower on first load, then provides ~10% prefill boost with a small reduction in token gen seen on larger models, and an increase on smaller models. However, it also enables NVFP4 input scale , which boosts model correctness. Initial performance results: llama bench on 5090: Perplexity/kld results against wiki2 test:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy