gpt oss 20b abliterated A refusal suppressed variant of openai/gpt oss 20b, produced with abliterix using direct weight editing , Expert Granular Abliteration (EGA) on the fused MoE expert weights, and MoE router suppression on the safety concentrated experts. Key results Metric Base gpt oss 20b This model Refusals on 100 held out harmful prompts (LLM judge) 97 / 100 6 / 100 KL divergence vs base (next token, benign) — 0.0098 Response length deviation vs base (benign) — 0.02 σ Hard prompt qualitative compliance (15 classic jailbreaks, EN+ZH) 0 / 15 15 / 15 The eval refusal counts come from an LLM judge ( google/gemini 3.1 flash lite preview via OpenRouter) instructed to label garbled / repetitive / incoherent output as a refusal — so models that "bypass" refusal by collapsing into gibberish get correctly counted as failures, not successes. A pre LLM rule based filter additionally catches dash runs, sentence loops, and low character diversity output before the judge is called. The 6/100 is a real, semantic compliance number, not keyword matching. The qualitative compliance row is a separate manual test: 15 classic hard prompts (10 EN + 5 ZH) covering lockpicking, phishing, meth synt…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy