Qwen3.6 27B IQ4 XS Pure with MTP GGUF This GGUF combines the Ununnilium IQ4 XS pure quantization of Qwen3.6 27B with the MTP (Multi Token Prediction) head extracted from the unsloth IQ4 XS MTP build, enabling native speculative decoding at a compact file size. File File Size Quantization qwen3.6 27b IQ4 XS pure with MTP IQ4.gguf ~13.57 GB Body: IQ4 XS, MTP: IQ4 NL/Q5 K/Q8 0 mix What's Inside Body (851 tensors) : IQ4 XS quantization via llama quantize pure from Ununnilium's build, using the unsloth imatrix for calibration MTP head (15 tensors) : Extracted from unsloth/Qwen3.6 27B MTP GGUF (IQ4 XS variant), preserving the original mixed quantization: attn q , attn k , attn output , ffn gate/up/down → IQ4 NL attn v → Q5 K nextn.eh proj → Q8 0 Norm tensors → F16 Why This Exists The Ununnilium pure GGUF strips MTP tensors to save space, but that means speculative decoding can't use the trained native draft head. This file grafts the MTP head back in, restoring native MTP speculative decoding while keeping the aggressive IQ4 XS body quantization for VRAM efficiency. Provenance Base model : Qwen/Qwen3.6 27B Body quantization : Ununnilium/Qwen3.6 27B IQ4 XS pure GGUF MTP head source : unsl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy