bottlecapai/ThinkingCap Qwen3.6 27B GGUF GGUF / llama.cpp quantizations of bottlecapai/ThinkingCap Qwen3.6 27B — capability of Qwen3.6 27B with 50% less thinking tokens on average, achieved by finetuning Qwen3.6 27B (Qwen Team, 2026) while preserving the original answer quality and style. ➡️ Full model description, evaluation results (multi seed, statistically tested), recommended sampling params, and citation: see the main model card at bottlecapai/ThinkingCap Qwen3.6 27B. About GGUF and quantization GGUF is a single file model format for running LLMs locally with llama.cpp and compatible runtimes (Ollama, LM Studio, …). The quantized variants below store weights at reduced precision — e.g. ≈4.7 bits per weight for Q4 K M instead of the 16 bit f16 source — cutting download size and memory severalfold at a small, measured quality cost. Files File Quant Size ThinkingCap Qwen3.6 27B Q4 K M.gguf Q4 K M 15.7 GB ThinkingCap Qwen3.6 27B Q6 K.gguf Q6 K 20.9 GB ThinkingCap Qwen3.6 27B Q8 0.gguf Q8 0 27.1 GB ThinkingCap Qwen3.6 27B f16.gguf f16 50.9 GB mmproj ThinkingCap Qwen3.6 27B f16.gguf mmproj (vision) 0.9 GB f16 is the unquantized source; Q8 0 is near lossless; Q6 K is a slightly smal…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy