Gemma 4 31B IT — Abliterated This is an abliterated version of google/gemma 4 31B it, created using Abliterix. This revision updates the model to trial 40 , the best configuration from the completed 60 trial Gemma 4 31B retraining run. Method Gemma 4's double norm architecture (4x RMSNorm per layer) and Per Layer Embeddings (PLE) make naive LoRA and hook based steering unreliable for this model family. This release uses direct weight editing : norm preserving orthogonal projection applied to the base model weights. Key techniques: Direct orthogonal projection on attention Q/K/V/O projections MLP down projection disabled for the selected run, improving stability for Gemma 4 31B Norm preserving row magnitude restoration , important for the double norm architecture float32 projection precision to avoid signal loss in high dimensional inner products Winsorized steering vectors (99.5th percentile) to reduce outlier activation influence Wider strength search range [1.0, 6.0] to explore beyond conservative low KL solutions vLLM in place evaluation during optimization, followed by a full HF safetensors export of the selected trial Evaluation Metric Value : Selected trial 40 Refusals (priva…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy