Qwen3.6 27B with MTP Up to 2.7× faster with MTP · 262K context on 48 GB · Fixed chat template Dense 27B model with vision, thinking, and tool use — self speculative decoding, \ configurable KV cache, fixed Jinja template (tool calls and thinking actually work in C++ runtimes), \ and a server with both OpenAI and Anthropic APIs. Start the server You need llama.cpp b9180 or newer (released 2026 05 16, includes MTP support). Install via Homebrew: Flag What it does Impact mmproj mmproj Qwen3.6 27B f16.gguf Vision encoder (text + image input) Multimodal support (+0.9 GB) spec type draft mtp spec draft n max 3 Multi Token Prediction (built into the model) Up to 2.7× faster generation c 262144 262K context window Full native context on 64 GB Mac fa off Disable Flash Attention 37–53% faster prefill on Apple Silicon at long context Sampling is set for coding tasks (temp 0.6, top p 0.95). Adjust m and c for your hardware — see the quant table below. For general chat, change to temp 0.7 top p 0.80 . Drop mmproj if you don't need vision. Optional flags 8 bit KV cache — halves KV memory at minor quality cost. Use when f16 KV doesn't give enough context: Custom chat template — override the embed…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy