Qwen3.6 27B abliterated v2 — GGUF GGUF quantizations of wangzhang/Qwen3.6 27B abliterated, a second pass refusal suppressed Qwen3.6 27B (10/100 refusals, 15/15 hard prompt compliance, cumulative KL ≈ 0.024 vs original Qwen/Qwen3.6 27B ). See the base model card for the full abliteration methodology and V1 → V4 sweep history. Files File Type Size VRAM (context 4k) Notes Qwen3.6 27B abliterated v2 F16.gguf F16 ~54 GB ~56 GB Lossless vs BF16 (small rounding on bfloat16 → float16 downcast). Reference quality. Qwen3.6 27B abliterated v2 Q8 0.gguf Q8 0 ~28 GB ~30 GB Near lossless. Recommended if you have 32+ GB VRAM or unified memory. Qwen3.6 27B abliterated v2 Q4 K M.gguf Q4 K M ~16 GB ~18 GB Recommended for 24 GB cards (3090/4090/A6000) and 24 GB Apple Silicon. K quant with mixed precision blocks (Q6 K for attn qkv + ffn down , Q4 K elsewhere). All three were produced from the same source BF16 safetensors checkpoint with llama.cpp convert hf to gguf.py (F16) and llama quantize (Q8 0, Q4 K M). No imatrix calibration — if you want better Q4 behaviour at minimal extra setup, run your own imatrix pass on a small held out prompt set and requantize from the F16 file here. llama.cpp version r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy