gemma 4 A4B 98e v7 coder it GGUF GGUF quantizations of ManniX ITA/gemma 4 A4B 98e v7 coder it, the loop fixed code prune of Gemma 4 26B A4B (128→98 experts/layer, ~20.8B). All quants made using imatrix with calibration data v5. The imatrix.dat used is included in this repo for reproducibility/audit; mmproj gemma4.gguf is the shared Gemma 4 SigLIP vision tower (untouched by pruning) for multimodal use. Quantizations — score, size & bits per weight Each published tier was scored on HumanEval+ (164) and MultiPL E 100 (llama.cpp, deployment sampler — min p ), as a quant vs quant comparison of how each tier holds the model's code ability. The table shows, per tier, the code score, the exact size, and the true bits per weight ( bpw = 8 × bytes ÷ 19,877,953,946 ). ⭐ marks a recommended pick . Tier Size bpw HE+ % MPE 100 % : : : : Q8 0 21.16 GB 8.52 91.46 91.00 Q6 K L 17.98 GB 7.24 91.46 89.67 Q6 K 17.81 GB 7.17 90.85 89.00 Q5 K L 15.25 GB 6.14 91.46 90.00 Q5 K M 15.07 GB 6.07 91.46 90.33 Q4 K L 13.42 GB 5.40 92.07 89.33 Q4 K M ⭐ 13.24 GB 5.33 92.68 89.00 Q4 K S 12.21 GB 4.91 92.68 89.33 IQ4 NL 11.42 GB 4.60 89.63 88.67 IQ4 XS 11.01 GB 4.43 91.46 88.33 Q3 K L 10.94 GB 4.40 90.85 89.00 Q3 K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy