Gemma 4 12B IT Abliterated — GGUF DuoNeural 2026 06 03 GGUF quantizations of DuoNeural/Gemma4 12B IT Abliterated — an abliterated Gemma 4 12B IT with the refusal direction surgically removed. Quantized with llama.cpp. Files File Size Recommended Use gemma4 12b abliterated Q4 K M.gguf ~7.5GB Best tradeoff — fits 12GB VRAM, excellent quality gemma4 12b abliterated Q5 K M.gguf ~8.5GB High quality, needs 12GB VRAM gemma4 12b abliterated Q8 0.gguf ~12.7GB Near lossless, needs 16GB VRAM Speed Benchmarks (A100 40GB, all layers GPU, llama bench) Quantization Size Prefill (tok/s) Generation (tok/s) Q4 K M 6.86 GiB 2,583 ± 139 78.3 ± 0.4 Q5 K M 7.95 GiB 2,455 ± 205 73.1 ± 0.2 Q8 0 11.78 GiB 2,573 ± 206 63.4 ± 0.3 Benchmarked on A100 40GB SXM4. ngl 99 (all layers to GPU). llama bench pp256/tg64. Usage (llama.cpp) Usage (Python via llama cpp python) Abliteration Details Base : google/gemma 4 12B it (48 layers, hidden=3840) Method : Orthogonal rank 1 projection (targeted mode: down proj + o proj, all 48 layers, α=0.3) Results : 5/7 harmful probes complied (71%) 6/6 benign probes preserved (100%) Mean KL Divergence (BF16→BF16, unbiased) : 0.0000 — zero measurable distribution shift on benign tex…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy