Qwen 3 4B Heretic (Abliterated) 💬 Community: Join the Abliterlitics Discord for discussion, model releases and support. An abliterated version of Qwen 3 4B created using Heretic v1.2.0. This model has reduced refusals while maintaining model quality, making it suitable as an uncensored text encoder for image generation models like Z Image and FLUX.2 Klein 4B. Model Details Base Model: Qwen/Qwen3 4B Abliteration Method: Heretic v1.2.0 Trials: 200 Trial Selected: Trial 96 Refusals: 3/100 (vs 100/100 original) KL Divergence: 0.0000 (zero measurable model damage) Files HuggingFace Format (for transformers, llama.cpp conversion) ComfyUI Format (for Z Image / FLUX.2 Klein 4B text encoder) GGUF Format (for llama.cpp and ComfyUI GGUF) Quant Size Notes F16 ~7.5GB Lossless reference Q8 0 ~4GB Excellent quality Q6 K ~3GB Very good quality Q5 K M ~2.7GB Good quality Q4 K M ~2.3GB Recommended balance Q3 K M ~1.9GB For low VRAM only NVFP4 Notes The NVFP4 (4 bit floating point, E2M1) variants use ComfyUI's native quantization format. They are ~3x smaller than bf16 and load natively in ComfyUI without any plugins. Blackwell GPUs (RTX 5090/5080, SM100+) can use native FP4 tensor cores for best per…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy