Gemma 4 12B OBLITERATED Zero refusal. Zero capability loss. First in the field. 0/842 refusals. 46/70 MMLU Pro (stock parity). Full coherence. The first abliterated model to achieve zero refusal with zero benchmark regression versus stock weights. Built with a novel 2 pass surgery pipeline developed by OBLITERATUS: 1. SOM Refusal Geometry Removal (Pass 1) — layers 12 21 2. ASPA Step Gradient Source Tethering (Pass 2) — layers 22 46 ⚠️ Research Context & Responsible Use This model exists for alignment research, red teaming, and safety evaluation. OBLITERATION is a weight surgery technique that studies how safety behaviors are geometrically encoded in transformer activation space. By precisely identifying and removing refusal directions, this research contributes to the scientific understanding of: How alignment is represented in model weights (mechanistic interpretability) How robust current safety training is against post training modification What the failure modes of RLHF/DPO based alignment are when adversaries have weight access This is the same class of research conducted by Arditi et al. ("Refusal in Language Models Is Mediated by a Single Direction", 2024), Zou et al. (HarmB…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy