Qwen3.6 27B i1 IQ4 KS GGUF This repository contains GGUF format weights for the Qwen3.6 27B model, quantized using the ik llama.cpp project. This model was specifically created to run on consumer GPUs with 16GB VRAM . By utilizing q4 0 KV cache quantization, it allows pushing the context length up to 105k tokens . ⚠️ Note: This model is designed exclusively for nVidia GPUs and is based on the advanced KS quants developed by ikawrakow from the ik llama.cpp repository. !!!UPDATE!!! I added a new quantization version which keeps the same file size but moves the model's weight from knowledge layers to logic layers. This results in a slight hit to PPL but should improve the model's performance on coding tasks. The new version is Qwen3.6 27B.i1 IQ4 KS attn qkv IQ4 KS . I also added a version with an MTP head. Quantization Details & Imatrix File Quantization Base: KS Quants (ik llama.cpp). Imatrix File Used: Provided by mradermacher. Other Tested Imatrix Files: bartowski – yielded significantly worse results. ubergarm – yielded comparable results. If you find or generate a better Imatrix file, please let me know in the Community tab! I have also included the script I used for the quantiza…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy