Gemma 4 12B Coder — SFT v5 + abliterated (GGUF) Uncensored gemma 4 12B coder for local, agentic tool use — GGUF quantizations for llama.cpp / Ollama. Run it: llama server hf tpls/gemma 4 12B coder fable5 composer2.5 v1 sft v5 abliterated GGUF:Q4 K M jinja (full commands below). ⚠️ Tool calling needs the recovery shim. The model emits gemma 4's native tool markup, which llama.cpp jinja under parses — wrap your endpoint with the tool shim (see Tool calling below) to get standard tool calls . 💡 Pick this to run locally with the best of both: SFT v5's tool calling and an uncensored model. Our KL guarded abliteration on top of SFT v5 — gate SHIM 8/8 . At a glance Type GGUF quantizations · llama.cpp / Ollama Techniques sft qlora → abliteration → imatrix quant → tool shim Tool calling ✅ 100% gate pass (recovery shim path) Status ✅ Active / supported Use llama server hf tpls/gemma 4 12B coder fable5 composer2.5 v1 sft v5 abliterated GGUF:Q4 K M jinja Use it Files Sizes and a one click loader are in the file browser / Quantizations widget above; the note says which quant to reach for. Quant Notes Q4 K M good default — fits 12 GB VRAM, best size/quality balance Q5 K M higher quality, ~9.5 G…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy