Qwen3.6 27B AWQ 4 bit (native) AWQ 4 bit quantization of Qwen3.6 27B (dense VL) with thinking + vision preserved, optimized for AMD RDNA4 (gfx1201) inference with SGLang. Model Details Base model Qwen/Qwen3.6 27B Architecture Qwen3.5 dense+DeltaNet hybrid + vision tower Parameters 27B Layers 48 (mixed full attention + DeltaNet linear attn) Context 262K (native) Modalities text + image + video (no audio) Quantization Native AWQ 4 bit, group size=128, fused Triton GEMM Calibration GPTQ via llmcompressor, 256 samples × 1024 tokens, thinking vision recipe; DeltaNet in proj a/b and vision tower kept BF16 Performance (2x AMD Radeon AI PRO R9700, TP=2) sglang.bench serving , single user, FP8 KV cache: Context TPOT (ms) tok/s : : : 128 41.5 24.1 8192 42.2 23.7 32768 54.5 18.3 65536 70.4 14.2 131072 102.4 9.8 Dense attention scales quadratically — the curve drops past 16K, unlike the 35B A3B MoE which stays flat. For long context coding/agent workloads on RDNA4, prefer the Qwen3.6 35B A3B AWQ (~21 tok/s flat through 131K). Notes Same calibration recipe + DeltaNet preservation as the 35B variant; differs only in dense vs MoE architecture. Native AWQ format (not compressed tensors) — preserve…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy