Qwen3.6 35B A3B AWQ Base model: Qwen/Qwen3.6 35B A3B This repo quantizes the model using data free quantization tool. (no calibration dataset was involved) 【Dependencies / Installation】 As of 2026 04 16 , make sure your system has cuda12.8 or cuda13.0 installed. Then, create a fresh Python environment (e.g. python3.12 venv) and run: vLLM Official Guide 【vLLM Startup Command】 Note: When launching with TP=8, include enable expert parallel ; otherwise the expert tensors wouldn’t be evenly sharded across GPU devices. 【Logs】 【Model Files】 File Size Last Updated 24GiB 2026 04 16 【Model Download】 【Overview】 Qwen3.6 35B A3B [!Note] This repository contains model weights and configuration files for the post trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Following the February release of the Qwen3.5 series, we're pleased to share the first open weight variant of Qwen3.6. Built on direct feedback from the community, Qwen3.6 prioritizes stability and real world utility, offering developers a more intuitive, responsive, and genuinely productive coding experience. Qwen3.6 Highlights This rele…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy