Read our How to Run Qwen3.5 MTP Guide! To run Qwen3.5 locally Read our Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. NEW: MTP speculative decoding for ~1.5 2x faster generation build llama.cpp from the MTP PR branch has been merged 16th May 2026! Note: llama.cpp renamed spec type mtp to spec type draft mtp on 2026 05 13 Set DGGML CUDA=OFF for CPU/Metal. np 1 and mmproj are not yet supported with MTP. Developer Role Support so Qwen3.5 can work in Codex, OpenCode and more! Tool calling improvements: Makes parsing nested objects to make tool calling succeed more. You can now also fine tune the model locally with Unsloth. Read our Qwen3.5 fine tuning guide here. Qwen3.5 0.8B [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intended use cases are prototyping, task specific fine tuning, and other research or development purposes. Over recent months, we have intensified our focus on developing foundation models that deliver e…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy