ThinkingCap Qwen3.6 27B INT4 AutoRound A fast INT4 / W4A16 AutoRound quantization of bottlecapai/ThinkingCap Qwen3.6 27B , prepared for local long context agent serving with vLLM. This is a practical first release: small enough to run comfortably on a 32 GB RTX 5090 class card, with Qwen3.6 thinking/tooling behavior preserved and Multi Token Prediction (MTP) tensors kept in BF16 for speculative decoding. A higher quality v2 quant is planned later. The likely direction is a more quality oriented recipe and/or stronger calibration, after we finish comparing the speed / context / quality tradeoffs. Credits Huge thanks to BottleCap AI for releasing ThinkingCap. Their checkpoint is the model here; this repo is only a community quantization of it. Also, personally: BottleCap are local friends/neighbours — their offices are about 200 meters from me, and I have known them for years — so this one is especially fun to package properly. Please give the original model and team credit: Base model: bottlecapai/ThinkingCap Qwen3.6 27B BottleCap AI: What is ThinkingCap? From the original BottleCap release: ThinkingCap keeps the capability and style of Qwen3.6 27B while using substantially fewer th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy