Gemma 4 31B it 4 bit AWQ Quantization (W4A16) This is a 4 bit AWQ quantization of Google's Gemma 4 31B it instruction tuned multimodal model, optimized for deployment with vLLM. Format note : This model was quantized using the AWQ algorithm via llm compressor and is saved in compressed tensors format. When loading with vLLM, use quantization compressed tensors (not quantization awq , which expects the AutoAWQ schema and will fail). Model Details Property Value Base Model google/gemma 4 31B it Quantization Algorithm AWQ (Activation aware Weight Quantization) Weight Scheme W4A16 ASYM (4 bit asymmetric weights, 16 bit activations) Weight Precision 4 bit int4 Activation Precision FP16/BF16 Group Size 128 Serialization Format compressed tensors (pack quantized) Quantization Library llm compressor (main branch, post 0.10.0.1) Architecture Gemma4ForConditionalGeneration Decoder Layers 60 Hidden Size 5376 Context Window 262K tokens (128K verified) Vision Tower SigLIP (27 layers, preserved in BF16 — NOT quantized) Quantized Components Text decoder only (vision tower + multimodal projector excluded) Hardware Requirements Verified context : 128K tokens on 48GB total VRAM (2× 24GB GPUs) and 25…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy