Qwen3.6 35B A3B NVFP4 MTP GGUF This repo contains two experimental NVFP4 GGUF quantizations of Qwen3.6 35B A3B for llama.cpp . This was quantized using my experimental advanced gguf quantizer tool. Both models were imatrix calibrated for the first time using a new custom dataset that I am evaluating. This repository contains two NVFP4 variants: Variant File Best for Notes TURBO Qwen3.6 35B A3B NVFP4 MTP TURBO.gguf Max speed More NVFP4. Lower quality metrics. HQ Qwen3.6 35B A3B NVFP4 MTP HQ.gguf Better quality More tensors promoted. Slightly slower. Quality & Speed Results All PPL/KLD results were measured against the same BF16 wikitest KLD base, and then compared to the official NVFP4 release by NVIDIA. Metric TURBO HQ NVIDIA NVFP4 : : : Size 18.56 GiB 18.64 GiB 22.20 GiB Mean PPL(Q) 6.987392 6.897796 7.014030 Mean PPL(Q) PPL(base) 0.268551 0.178955 — Mean PPL ratio 1.039970 1.026635 1.043935 Mean ln(PPL ratio) 0.039192 0.026286 — Mean KLD 0.063228 0.050759 0.066331 99.9% KLD 1.924147 1.565143 1.560988 99.0% KLD 0.598519 0.488387 0.495896 95.0% KLD 0.221030 0.178889 0.207580 Max KLD 11.946571 10.093911 6.972712 Same top p 89.023% 90.255% 87.608% Top flip weight 0.012068 0.009575 —…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy