Qwen 3.6 27B — MXFP4 (MLX) Open Compute Project MXFP4 quantization of Alibaba's hybrid linear/full attention dense 27B VL model, with the vision tower preserved. Model Details Property Value Base model Qwen/Qwen3.6 27B Parameters 27.32 B, dense (no MoE) Architecture qwen3 5 — 64 decoder layers: 48 Gated DeltaNet (linear attn) + 16 full attention with swish output gate Quantization OCP MXFP4 (E2M1 + shared E8M0 scale) at block 32 Package size on disk 14 GB across 3 shards Bits per weight 4.449 vs BF16 source 52 GB → 14 GB, 3.7× compression Context (position embeddings) 262,144 native; upstream card reports up to ~1 M with YaRN scaling Vision tower 27 layer ViT (hidden 1152, patch 16), MXFP4 quantized Chat format Qwen im start/im end, unified thinking toggle Quantization details Category Bits Group Notes Dense FFN ( mlp.gate proj , mlp.up proj , mlp.down proj ) 4 (MXFP4) 32 Bulk of parameters Full attention projections ( q proj , k proj , v proj , o proj ) 4 (MXFP4) 32 q proj is fused with a swish output gate (output split 50/50 queries/gate) Linear attention projections ( in proj qkv , in proj z , in proj b , in proj a , out proj ) 4 (MXFP4) 32 Embedding ( embed tokens…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy