gemma 4 12B it heretic GGUF GGUF quantizations of igorls/gemma 4 12B it heretic, a fully automatic decensored ("abliterated") version of google/gemma 4 12B it produced with Heretic. The decensored model has 0/100 genuine refusals on harmful prompts at a KL divergence of only 0.0284 from the original model — censorship removed with minimal loss of capability. ⚠️ Use non thinking mode for best results Gemma 4 is a hybrid thinking model, and the abliteration targets the direct (non thinking) response — which is also Gemma 4's own default. For roleplay, creative writing, and the most reliable uncensored output, run with thinking disabled. In thinking mode the model produces good output too, but the chain of thought consumes the token budget and can leave the final answer truncated. Runtime How to disable thinking : : Ollama (CLI) /set nothink in the session Ollama (API) add "think": false to the request body llama.cpp omit jinja , or use a prompt that closes the thought block transformers already non thinking by default ( enable thinking=False ) If you do use thinking mode, set a large num predict / num ctx so the answer isn't cut off by the reasoning block. Files File Quant Size Notes…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy