gemma 4 26B A4B it uncensored Uncensored version of google/gemma 4 26B A4B it with refusal behavior removed. Results Before After Refusals (mlabonne, 100 prompts) 98/100 1/100 effective (3 flagged, 2 refusal then comply) Refusals (cross dataset, 686 prompts) — 5/686 (0.7%) KL Divergence 0 (baseline) 0.09 Quality (harmless response length ratio) 1.0 ~1.01 (no degradation) Cross Dataset Validation Tested against 4 independent prompt datasets to verify generalization: Dataset Prompts Refusals JailbreakBench 100 1/100 tulu harmbench 320 1/320 NousResearch/RefusalDataset 166 0/166 mlabonne/harmful behaviors 100 3/100 Total 686 5/686 (0.7%) Every flagged refusal was manually audited. Most are "refusal then comply" false positives where the model adds an AI identity disclaimer then answers the question anyway. Method Norm preserving biprojected abliteration on the dense pathway (o proj + shared mlp.down proj), plus Expert Granular Abliteration (EGA) on all 128 MoE expert down proj slices per layer. EGA (OBLITERATUS) hooks the MoE routers during probing to compute per expert routing weights for harmful vs harmless prompts, then applies norm preserving projection (grimjim) to each expert in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy