Gemma 4 12B (QAT Q4 0 unquantized) — Heretic decensored — GGUF GGUF quantizations of igorls/gemma 4 12B it qat q4 0 unquantized heretic, a decensored ("abliterated") version of Google's Gemma 4 12B QAT Q4 0 unquantized checkpoint, produced with Heretic v1.3.0. Quant Size Notes : : : : Q4 0 6.5 GB Recommended. QAT matched: the base was quantization aware trained for the Q4 0 grid, so this 4 bit quant best preserves the original quality Q4 K M 6.9 GB Standard K quant 4 bit alternative Q8 0 11.8 GB Near lossless [!NOTE] This model derives from Google's QAT Q4 0 checkpoint. Quantization aware training was calibrated specifically for the Q4 0 grid, so Q4 0 is the quant that realizes the QAT advantage (≈bf16 quality at 4 bit). Q4 K M is a normal K quant of the same weights and does not specifically leverage QAT. v1.1 — thinking mode fix Gemma 4 is a thinking model: its refusal decision forms inside the chain of thought. The first release was abliterated/evaluated with thinking disabled, so it still refused once thinking was on (the default in Ollama / llama.cpp). v1.1 decensors the model with thinking enabled — the way it's actually used. Performance Metric This model (v1.1) Original : :…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy