Read our How to Run Qwen3.5 MTP Guide! To run Qwen3.5 locally Read our Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. NEW: MTP speculative decoding for ~1.5 2x faster generation build llama.cpp from the MTP PR branch has been merged 16th May 2026! Note: llama.cpp renamed spec type mtp to spec type draft mtp on 2026 05 13 Set DGGML CUDA=OFF for CPU/Metal. np 1 and mmproj are not yet supported with MTP. Developer Role Support so Qwen3.5 can work in Codex, OpenCode and more! Tool calling improvements: Makes parsing nested objects to make tool calling succeed more. Disable thinking via chat template kwargs '{"enable thinking":false}' . Read our guide. Mar 5 'Final' Update: All GGUFs now use our new imatrix data. See some improvements in chat, coding, long context, and tool calling use cases. GGUFs now updated with an improved quantization algorithm. Rest of variants like Q8 0, Q4 K M, BF16 are now uploaded. See our new benchmarks for 122B A10B here. For Qwen3.5 35B A3B, we primarily reduced the maximum KLD: Feb 27 Update: GGUFs Refreshed + Tool calling fixes + Benchmarks Qwen3.5 is now updated with improved tool calling & coding performance! S…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy