ThinkingCap Qwen3.6 27B — NVFP4 NVFP4 (4 bit) quantization of bottlecapai/ThinkingCap Qwen3.6 27B — the token efficient finetune of Qwen/Qwen3.6 27B that produces ~50% fewer thinking tokens on average (up to ~90% on individual examples) while preserving answer quality. This quant makes ThinkingCap practical on memory bandwidth limited Blackwell hardware: weights shrink from ~55 GB (BF16) to ~20.6 GB, which on an NVIDIA DGX Spark (GB10, 273 GB/s unified memory) takes single stream decode from 4.6 tok/s to a measured 26.0 tok/s (with the preserved MTP speculative decoding, n=3) — a 5.7× speedup — or 62.5–64.1 tok/s (up to 13.9×) with an external DFlash drafter (n=12, BF16 KV; see version compatibility notes below). Quantized by MoroSystems — on a DGX Spark itself. What's inside Weights: NVFP4 (FP4 with FP8 block scales, group size 16), weights + activations (W4A4), produced with NVIDIA TensorRT Model Optimizer 0.45.0 ( quant method: modelopt ). Kept in BF16 for quality: lm head , every linear attn.conv1d (gated deltanet short conv), the full vision tower ( model.visual ), and the MTP head ( mtp ) — so native qwen3 next mtp speculative decoding still works . Also kept in BF16 (more co…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy