Qwen3.5 9B abliterated Unrestricted version of Qwen/Qwen3.5 9B, created with Abliterix — automated LLM abliteration via orthogonalized steering and Bayesian optimization. Highlights Metric Value Refusal rate 2/200 (1%) KL divergence 0.0105 Optimization trials 50 Strong balance of capability and efficiency at 9B parameters: 1% refusal rate with practical VRAM requirements. How It Works Abliterix removes safety refusal behavior while preserving model capabilities: 1. Refusal direction extraction — 800 harmful + 800 benign prompts reveal per layer refusal activation patterns 2. Orthogonal projection — isolates the refusal signal by projecting out components aligned with normal responses, reducing refusals by 67% vs. raw abliteration 3. LoRA based abliteration — rank 1 modifications to attention and MLP weights, captured as lightweight adapters (not destructive edits) 4. Bayesian optimization — Optuna TPE searches kernel shape, fractional direction index, and per component strength across 50 trials to find the Pareto optimal balance of low refusals and low KL divergence All Abliterix Models Model Refusals KL Divergence Trials Qwen3.5 122B A10B abliterated 1/200 (0.5%) 0.0115 25 Qwen3.5…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy