Qwen3.5 122B A10B GPTQ 4 bit GPTQ 4 bit quantization of Qwen/Qwen3.5 122B A10B, a 122B parameter Mixture of Experts (MoE) multimodal model with ~10B activated parameters per token. Includes full vision encoder and MTP (Multi Token Prediction) module for image understanding and speculative decoding support. Model Overview Architecture : Qwen3 5MoeForConditionalGeneration (multimodal: text + vision) Total parameters : ~122B Activated parameters : ~10B per token (8 of 256 experts selected per token) Layers : 48 (36 linear attention + 12 full attention, repeating 3:1 pattern) Experts : 256 per layer + 1 shared expert per layer Context length : 262,144 tokens Vision encoder : 27 block ViT (1152 hidden, 16x16 patches), BF16 MTP module : 1 layer speculative decoding head, BF16 Quantization Details All 36,864 MoE expert modules (256 experts x 3 projections x 48 layers) are quantized to INT4 using GPTQ. Non expert modules (including the full vision encoder and MTP module) remain at BF16/FP16 for quality preservation. Component Precision Notes MoE experts ( gate proj , up proj , down proj ) INT4 (GPTQ) 36,864 modules quantized Full attention ( q proj , k proj , v proj , o proj ) FP16 Every 4…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy