Qwen3.6 27B GPTQ Int4 GPTQ Int4 quantization of Qwen/Qwen3.6 27B, produced on consumer multi GPU hardware (4× RTX 3060 12GB) using Python 3.13t free threading. This release uses uniform group size=32 for full vLLM/SGLang compatibility and ships MTP (Multi Token Prediction) speculative decoding weights verified working on vLLM 0.21.0. Quality Metric Value GPTQ success rate 100% RTN fallback rate 0% Loss mean 1.51e 04 Loss max 8.73e 04 Total modules 400 What's quantized vs kept bf16 Quantized (int4, g32 uniform): Self attention: q proj , k proj , v proj , o proj Linear attention (GDN/Mamba hybrid): in proj qkv , in proj z , out proj MLP: gate proj , up proj , down proj Kept bf16 (per Qwen3.6 FP8 recipe): Linear attention state dynamics: A log , conv1d , dt bias , in proj b , in proj a , in proj ba , linear attn.norm Attention norms: q norm , k norm Layer norms: input layernorm , post attention layernorm Multi token prediction head: mtp. (15 keys) Vision encoder: model.visual. (333 keys — ViT blocks for image input) Embeddings ( model.language model.embed tokens ) and lm head Calibration recipe Domain mixed calibration set: Source Samples Purpose allenai/c4 102 General English text al…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy