gemma 4 31b he1 it GGUF GGUF quantizations of ManniX ITA/gemma 4 31b he1 it. All quants made using imatrix with calibration data v5. Available Quantizations Quantization Status Q8 0 30.39 GB Q6 K L 23.79 GB Q6 K pending Q5 K L pending Q5 K M pending Q5 K S 19.85 GB Q4 K L pending Q4 K M pending Q4 1 18.14 GB Q4 K S pending Q4 0 pending IQ4 NL 16.44 GB IQ4 XS 15.59 GB Q3 K XL 14.56 GB IQ3 M 13.43 GB Q3 K L 15.49 GB Q3 K M 14.24 GB Q3 K S 12.82 GB IQ3 XS 12.17 GB IQ3 XXS 11.25 GB Q2 K L 11.42 GB IQ2 M 10.17 GB IQ2 S 9.46 GB IQ2 XS 8.88 GB How to Use With llama.cpp: With ollama (requires Modelfile or HF direct load). Original Model Card gemma 4 31B he1 it — v2 (2026 05 17) Partial head prune of google/gemma 4 31B it : 12.5% of Q heads removed in the first 4 sliding attention layers (L0–L3), with an lstsq heal of those layers' O projections. L4–L59 are byte identical to the base. The model matches base model quality on HumanEval and MBPP under chat completions evaluation; the early layer perturbation is what the local rebuild captured, and the gains come from the heal solution found on those 4 layers. The build is fully reproducible via OmniMergeKit ( recipes/gemma4 31b/prune local hea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy