Gemma 4 26B A4B MoE IT QAT Assistant — INT4 W4A16 Compressed Tensors (Marlin) This is the MTP draft/assistant model for Gemma 4 26B A4B MoE, quantized from google/gemma 4 26B A4B it qat q4 0 unquantized assistant (BF16) to INT4 W4A16 compressed tensors format with group size=128 , optimized for Marlin kernel acceleration. It is designed to be used alongside NeoChen1024/gemma 4 26B A4B it qat W4A16 as the main model with SGLang's Frozen KV MTP speculative decoding. Why This Model? Using the original BF16 assistant model with MTP speculative decoding causes a severe short sequence regression on Ampere GPUs when the main model uses Marlin INT4 kernels. This INT4 quantized draft model uses the same Marlin kernels, eliminating the regression and delivering a significant speedup: Config 100 tok 500 tok Main model only (no MTP) 145 tok/s 149 tok/s MTP + this INT4 draft ~183 tok/s ~144 tok/s 26% faster on short sequences vs no MTP baseline. Comparable on long sequences. Benchmarks on NVIDIA RTX 3090 (24 GB) with SGLang 0.5.12, mem fraction static=0.85 , cuda graph max bs=1 , context length=4096 , speculative num steps=3 , speculative num draft tokens=4 . How It Was Created Standard quantiz…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy