Qwen3.5 397B A17B AWQ Base model: Qwen/Qwen3.5 397B A17B This repo quantizes the model using data free quantization (no calibration dataset required). 【Dependencies / Installation】 As of 2026 02 25 , make sure your system has cuda12.8 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: vLLM Official Guide 【vLLM Startup Command】 Note: When launching with TP=8, include enable expert parallel ; otherwise the expert tensors wouldn’t be evenly sharded across GPU devices. 【Logs】 【Model Files】 File Size Last Updated 228GiB 2026 02 25 【Model Download】 【Overview】 Qwen3.5 397B A17B [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. [!Tip] For users seeking managed, scalable inference without infrastructure maintenance, the official Qwen API service is provided by Alibaba Cloud Model Studio. In particular, Qwen3.5 Plus is the hosted version corresponding to Qwen3.5 397B A17B with more production features, e.g., 1M context length by default, official built in tools, and adaptive tool use. For mor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy