Huihui ThinkingCap Qwen3.6 27B abliterated NVFP4 NVFP4 (W4A4) quantization of huihui ai/Huihui ThinkingCap Qwen3.6 27B abliterated — Huihui's abliterated (refusal removed / uncensored) finetune of bottlecapai/ThinkingCap Qwen3.6 27B , itself a token efficient reasoning fine tune of Qwen3.6 27B. Produced with llm compressor → compressed tensors , with the native MTP speculative decode head preserved (bf16) and the Qwen3 VL vision tower preserved (bf16). Why this pairing is nice. You keep ThinkingCap's short token efficiency (the base cuts reasoning length by ~46 % vs Qwen3.6 27B) and Huihui's abliteration (refusal directions removed), then NVFP4 + the MTP draft cut the cost of every token. Fewer thinking tokens × faster tokens × no refusal detours = a snappy, compliant local reasoner. Abliteration can shift behavior on some prompts — evaluate for your use case. 20.6 GB on disk (down from ~55.6 GB bf16). Serves on stock vLLM 0.21+ — no quantization flag needed (auto detected). Architecture Qwen3 5ForConditionalGeneration (model type qwen3 5 ), dense 27.4 B : Hybrid attention — Gated DeltaNet (linear) + full attention layers, hidden 5120, 262 K native context. Vision — Qwen3 VL ViT, k…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy