gemma 4 12b it uncensored — GGUF GGUF quantizations of zaakirio/gemma 4 12b it uncensored , a decensored (Heretic abliterated) version of google/gemma 4 12B it . These files run with llama.cpp . Gemma 4 12B (the "Unified" release, June 2026) is an encoder free unified multimodal model: text, image, audio, and video all project straight into a single decoder only transformer, with a context window of up to 256K tokens . Abliteration touches only the language weights, so those capabilities carry over unchanged. Files Filenames follow gemma 4 12b it uncensored .gguf . Quant Size Notes Q2 K 4.50 GB Smallest; lowest quality. Very tight memory only. Q3 K M 5.67 GB Small; usable on low RAM. Q4 K S 6.54 GB Compact 4 bit. Q4 K M 6.87 GB Recommended — best size/quality balance. Q5 K M 7.96 GB Higher quality, slightly larger. Q6 K 9.11 GB Near lossless. Q8 0 11.80 GB Effectively lossless vs the BF16 source. f16 22.20 GB Full precision; reference / re quantizing. Not sure which to pick? Start with Q4 K M. Go up to Q5/Q6/Q8 if you have the memory and want maximum fidelity; drop to Q3/Q2 only if you're memory constrained. Multimodal projector (for image/audio input — see Multimodal): File Size N…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy