Qwen3.6 35B A3B INT8 W8A8 INT8 ( W8A8 ) quantization of Qwen/Qwen3.6 35B A3B — a hybrid Mixture of Experts model (256 experts, top 8, ~3B active) with GatedDeltaNet linear attention, a vision tower and an MTP head. Quantization Symmetric INT8 weights (per channel, MSE observer) + INT8 dynamic per token activations (llm compressor). int8 is robust, so no AWQ/SmoothQuant smoothing is needed. Quantized: the MoE expert FFNs only (the bulk of the weights). Kept bf16 (quality sensitive): self attn , the router ( mlp.gate ), shared expert , GatedDeltaNet ( linear attn ), lm head , embeddings, vision tower, MTP head. Format: compressed tensors ( int quantized ). Full recipe: recipe.yaml . INT8 W8A8 is near lossless; for the smallest footprint use the INT4 W4A16 variant (GSM8K 96.8% / MMLU Pro 80.2%). Usage (vLLM) Served via vLLM's INT8 MoE path (works on Ampere sm 80 / sm 86). Quantized with vllm ampere optimized/quantize.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy