Gemma 4 31B it abliterated heretic AWQ W4A16 AWQ W4A16 (group size 128, symmetric) quantization of trohrbaugh/gemma 4 31b it heretic ara — a Heretic ARA abliterated derivative of google/gemma 4 31b it. ⚠️ Decensored model. Safety guardrails have been deliberately removed. Research and experimentation only. See full disclaimer below. Quantization Details Parameter Value : : : Method AWQ (Activation aware Weight Quantization) Scheme W4A16 (symmetric) Weight Bits 4 Activation Bits 16 Group Size 128 Format compressed tensors Calibration Dataset HuggingFaceH4/ultrachat 200k Calibration Samples 256 Max Sequence Length 2048 Vision Tower Unquantized (full precision) LM Head Unquantized (full precision) Compatible Inference Engine vLLM ( vllm/vllm openai:gemma4 ) Quantization Notes All multimodal paths kept full precision : Vision tower, audio tower, video tower, multi modal projector, and all modality specific embedding and projection layers are excluded from quantization. Only language model linear layers (attention Q/K/V/O and MLP gate/up/down) are quantized. LM head unquantized : Standard practice to preserve output token distribution quality at negligible size cost. v proj → o proj smo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy