Qwen3 Coder REAP 25B A3B AWQ 4 bit AWQ 4 bit quantization of Cerebras Qwen3 Coder REAP 25B A3B — a REAP pruned (arxiv:2510.13999) variant of Qwen3 Coder 30B A3B Instruct — calibrated with thinking + code data, optimized for AMD RDNA4 (gfx1201) inference with SGLang. Model Details Base model cerebras/Qwen3 Coder REAP 25B A3B (REAP prune of Qwen3 Coder 30B A3B Instruct) Architecture Qwen3 MoE (96 experts post REAP, top 8) Parameters ~25B total / 3B active Pruning method REAP (Router aware Expert pruning, 25% drop) — distinct from REAM (expert merging) Layers 48 Context 131K (tested), 256K supported by base Quantization Native AWQ 4 bit, group size=128, fused Triton GEMM Calibration GPTQ via llmcompressor, 256 samples × 1024 tokens, code thinking mix (AM Thinking v1, NuminaMath CoT, ultrachat); ignore= lm head, mlp.gate, shared expert. Performance (2x AMD Radeon AI PRO R9700, TP=2, fp8 KV) sglang.bench serving , single user, FP8 KV cache, disable cuda graph : Context TPOT (ms) tok/s : : : 128 43.6 22.9 1024 43.7 22.9 8192 44.1 22.7 32768 44.2 22.6 65536 45.5 22.0 131072 45.6 21.9 Flat ~22.5 tok/s decode across the full 131K range — A3B MoE stays bandwidth bound, no attention scaling c…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy