Qwen3.6 35B A3B GPTQ Int4 GPTQ Int4 quantization of Qwen/Qwen3.6 35B A3B, produced on consumer multi GPU hardware (4× RTX 3060 12GB) using Python 3.13t free threading. This v2 release ships MTP (Multi Token Prediction) speculative decoding weights verified working on both vLLM 0.19.1 and SGLang 0.5.10. Quality Metric Value GPTQ success rate 97.42% RTN fallback rate 2.58% Loss mean 1.38e 04 Loss median 9.29e 05 Loss max 2.14e 03 Total modules 30,720 Perplexity (wikitext 2 raw v1) 6.1846 (~97.9% BF16 retention) Model specs Property Value Base model Qwen3.6 35B A3B (MoE, 35B total / 3B active) Architecture Qwen3 5MoeForConditionalGeneration (vision + text) Experts 256 (top 8 routing per token) Hidden layers 40 Context length 262,144 tokens Quantization GPTQ v2, 4 bit, group size=128, symmetric Quantized size 24.4 GB (incl. MTP weights) KV cache support fp16, bf16, fp8 e4m3 (storage only on Ampere) MTP head Included (BF16, 785 keys, split per expert format) What's quantized vs kept bf16 Quantized (int4): All MoE expert weights ( mlp.experts. ) across layers 0–39 Kept bf16 (per Qwen3.6 recipe): Attention layers ( .attn. ) MoE routers ( .mlp.gate ) Shared experts ( .shared expert. ) Mult…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy