Qwen3.6 27B Pure Quantized Update: Previously I named this model Q4 K M pure , which is misleading. So the models has been renamed to 4.5bpw pure instead, to reflect the actual weight distribution. Weight Type Distribution KLD Comparison Overall comparison against Q3 and Q4 quants: More detailed on Q4 quants: Token speed On RTX 5060 Ti 16GB: Version Prompt Processing Token Generation MTP 195 tok/s 40 tok/s Non MTP 715 tok/s 24 tok/s MTP: Non MTP: Similar to Ununnilium/Qwen3.6 27B IQ4 XS pure GGUF, this is the quantization of Qwen3.6 27B GGUF using the pure param. For the Q4 K M version (both MTP and non MTP). The GGUF was generated with the Unsloth imatrix like this: pure means to use the same Q4 K M for nearly all tensors. (Original model card) Qwen3.6 27B [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Following the February release of the Qwen3.5 series, we're pleased to share the first open weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability an…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy