Qwen3.6 35B A3B — Abliterated V2 This is V2 of the abliterated (uncensored) Qwen/Qwen3.6 35B A3B, created using Abliterix. V2 improves on V1 by adding projected abliteration (grimjim 2025), outlier winsorization , 2× training data , and a larger TPE search budget — cutting the refusal rate from 7/100 to 4/100 under the same LLM judge evaluation. V1 vs V2 at a glance Metric V1 V2 (this model) Change Refusals (LLM judge, 100 eval prompts) 7/100 4/100 −43% Attack success rate 93% 96% +3 pt KL divergence from base 0.0189 0.0421 +0.023 Optimization trials completed 24/50 33/50 TPE explored more Training prompts 400 800 2× more data Eval prompts 100 100 (unchanged for fair A/B) V2 trades a small KL increase (still well under 0.1, no perceptible coherence loss) for a meaningful refusal rate improvement and a more robust steering vector trained on 2× the data. Method Qwen3.6 35B A3B is a Mixture of Experts model (256 routed experts, 8 active per token, 35B total / 3B active parameters) sharing identical architecture with Qwen3.5 35B A3B. Standard LoRA based abliteration is effective on this architecture (unlike Gemma 4's double norm design which requires direct weight editing). V2 inherits…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy