Qwen3.6 35B A3B abliterated v4 GGUF GGUF quantizations of Bahushruth/Qwen3.6 35B A3B abliterated v4 for llama.cpp, Ollama, LM Studio, KoboldCPP, and other GGUF compatible runtimes. Uncensored model — refusal behavior removed via norm preserving abliteration (0% refusal, full capability preservation). See the base model card for method details. Blog post: Abliteration: Uncensoring LLMs via Weight Surgery Quantizations All standard quants are compatible with Ollama, LM Studio, KoboldCPP, and llama.cpp out of the box — no special flags needed. File Size RAM Required Notes ... BF16.gguf 69 GB 80+ GB Full precision ... Q8 0.gguf 37 GB 48+ GB Near lossless ... Q6 K.gguf 29 GB 40+ GB Very high quality ... Q5 K M.gguf 25 GB 32+ GB Recommended for 48GB systems ... Q4 K M.gguf 21 GB 24+ GB Good quality, fits 24GB ... IQ4 XS.gguf 19 GB 24+ GB High quality 4 bit (imatrix) ... IQ4 NL.gguf 20 GB 24+ GB Non linear 4 bit (imatrix) ... Q3 K M.gguf 17 GB 20+ GB Good for 16GB VRAM GPUs ... IQ3 M.gguf 15 GB 20+ GB High quality 3 bit (imatrix) ... IQ3 XXS.gguf 14 GB 16+ GB Smallest usable 3 bit (imatrix) ... Q2 K.gguf 13 GB 16+ GB 2 bit, quality tradeoffs ... IQ2 M.gguf 12 GB 16+ GB Smallest, significa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy