Qwen3.6 27B Quark W8A8 INT8 W8A8 INT8 quantized version of Qwen/Qwen3.6 27B using AMD Quark. Model Details Base Model Qwen/Qwen3.6 27B Architecture Qwen3 5ForConditionalGeneration (hybrid attention + ViT) Parameters 27B language tower (quantized) + 27 layer ViT (BF16, unquantized) Layers 64 hybrid (16 full attention + 48 linear attention GatedDeltaNet) + 1 MTP head Quantization W8A8 INT8 (per channel weight + per token dynamic activation) Quantizer AMD Quark 0.11.1 ( pack method='reorder' , vLLM native key naming) Model Size ~29 GB (single safetensors) Original Size ~52 GB (BF16) Compression ~1.8x size reduction Quantization Scheme Component dtype Granularity Mode Linear weight (text decoder) INT8 per channel ( ch axis=0 ) symmetric, static Linear activation INT8 per token ( ch axis=1 ) symmetric, dynamic lm head BF16 unquantized embed tokens BF16 unquantized Vision tower (27 ViT blocks) BF16 unquantized MTP head ( mtp ) BF16 unquantized Accuracy GSM8K full 1319 question test split (vLLM, temperature=0 , concurrency=16 , max tokens=1024 , chat template kwargs.enable thinking=false ): Model Accuracy Correct Qwen/Qwen3.6 27B (BF16 baseline) 96.74% 1276 / 1319 This model (Quark W8A8 I…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy