Qwen3.6 35B A3B GPTQ 4 bit GPTQ 4 bit quantization of Qwen/Qwen3.6 35B A3B, a 35B parameter Mixture of Experts (MoE) multimodal model with 3B activated parameters per token. Includes full vision encoder and MTP (Multi Token Prediction) module for image understanding and speculative decoding support. Model Overview Architecture : Qwen3 5MoeForConditionalGeneration (multimodal: text + vision; same architecture as Qwen3.5) Total parameters : ~35B Activated parameters : ~3B per token (8 of 256 experts selected per token) Layers : 40 (30 linear attention + 10 full attention, repeating 3:1 pattern) Experts : 256 per layer + 1 shared expert per layer Context length : 262,144 tokens Vision encoder : 27 block ViT (1152 hidden, 16x16 patches), BF16 MTP module : 1 layer speculative decoding head, BF16 Quantization Details All 30,720 MoE expert modules (256 experts x 3 projections x 40 layers) are quantized to INT4 using GPTQ. Non expert modules (including the full vision encoder and MTP module) remain at BF16/FP16 for quality preservation. Component Precision Notes MoE experts ( gate proj , up proj , down proj ) INT4 (GPTQ) 30,720 modules quantized Full attention ( q proj , k proj , v proj ,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy