Qwen3.6 27B Omnimerge v4 MTP GGUF GGUF quantizations of ManniX ITA/Qwen3.6 27B Omnimerge v4 with the MTP (Multi Token Prediction) head retained for self speculative decoding on llama.cpp mainline (PR 22673, merged 2026 05 16) and later. Companion to the standard decode release at ManniX ITA/Qwen3.6 27B Omnimerge v4 GGUF . The two repos contain identical merged weights — this one keeps the additional mtp. tensors that convert hf to gguf.py remaps to blk.{num hidden layers}. per llama.cpp PR 22673 ("llama + spec: MTP Support", merged 2026 05 16), so spec type draft mtp works out of the box. All quants made with imatrix using bartowski's calibration datav5; imatrix.dat archived alongside the quants for reproducibility/audit. Available Quantizations Quantization Size (GiB) Notes F16 50.90 full precision reference Q8 0 27.05 Q6 K 20.89 recommended speed/quality balance Q5 K M 18.19 Q4 K M 15.66 IQ4 XS 14.26 IQ3 M 11.89 IQ2 M 9.54 MTP head forced to Q4 K — see note below IQ2 M MTP head override. The 7 MTP head tensors ( blk.64.attn {k,q,v,output}.weight , blk.64.ffn {down,gate,up}.weight , blk.64.nextn.eh proj.weight ) are overridden to Q4 K instead of the K mix's default IQ2 S/IQ3 S. Re…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy