To run Qwen3.5 locally Read our Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. Disable thinking via chat template kwargs '{"enable thinking":false}' . Read our guide. Mar 6 Update: All GGUFs now use our new imatrix data. See some improvements in chat, coding, long context, and tool calling use cases. GGUFs now updated with an improved quantization algorithm. See our new benchmarks for Qwen3.5 here. Qwen3.5 397B A17B [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. [!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Alibaba Cloud Model Studio. In particular, Qwen3.5 Plus is the hosted version corresponding to Qwen3.5 397B A17B with more production features, e.g., 1M context length by default, official built in tools, and adaptive tool use. For more information, please refer to the User Guide. Over recent months, we have intensified our focus on developing foundation models that deliver excep…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy