ThinkingCap Qwen3.6 27B — MTP GGUF (Blackwell / speculative decoding edition) GGUF quantizations of BottleCapAI/ThinkingCap Qwen3.6 27B — their "brevity finetune" of Qwen3.6 27B that keeps full accuracy with ~46% fewer thinking tokens — repackaged for the two things stock GGUFs don't give you: NVFP4 — native FP4 for Blackwell (RTX 50 series / RTX PRO 6000) tensor cores. MTP baked into every quant — the Multi Token Prediction draft head travels inside each file, so you get speculative decoding for free ( spec type draft mtp ), no second model to wire up. A full low bit ladder — IQ2→Q8 0 + a lossless bf16 master, so it fits everything from a 24GB card to a workstation. Same weights as upstream. Strictly more ways to run them, faster. Files quant size notes NVFP4 MTP 18.2 GB ← Blackwell FP4 tensor cores + MTP. The one to grab on RTX 50xx / PRO 6000. bf16 MTP 54.7 GB lossless master (exact bf16, not a lossy f16 re map) + MTP Q8 0 MTP 29.0 GB near lossless reference + MTP Q6 K MTP ~22 GB + MTP Q5 K M MTP ~19 GB + MTP Q4 K M MTP ~17 GB + MTP (the common daily driver size — and here it carries the draft head) IQ4 XS MTP ~14 GB imatrix low bit + MTP [fast follow] Q3 K M MTP 13.5 GB imatrix…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy