Qwen3.6 35B A3B Quark W8A8 INT8 W8A8 INT8 quantized version of Qwen/Qwen3.6 35B A3B produced with AMD Quark. Model Details Base Model Qwen/Qwen3.6 35B A3B Architecture Qwen3 5MoeForConditionalGeneration (multimodal: ViT vision + text MoE + MTP head) Parameters 35B total / 3B activated per token (256 experts, top 8) + 27 block ViT (BF16) Quantization W8A8 INT8 — per channel weight + per token dynamic activation Quantizer AMD Quark 0.11.1 ( pack method='order' , weight format='real quantized' ) Model Size ~35 GB (7 shards of ~5 GB) Original Size ~67 GB (BF16, 26 shards) Compression ~1.93× size reduction Quantization Scheme Component dtype Granularity Mode Language attention ( q/k/v/o proj , linear attn. ) INT8 per channel weight (axis=0) weight static Language MoE experts (256 × gate/up/down proj × 40) INT8 per channel weight (axis=0) weight static shared expert ( gate/up/down proj ) INT8 per channel weight (axis=0) weight static All activations above INT8 per token (axis=1) dynamic lm head BF16 — unquantized embed tokens BF16 — unquantized MoE router ( mlp.gate ) — top k gate BF16 — unquantized shared expert gate BF16 — unquantized visual. (27 block ViT + merger) BF16 — unquantized…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy