Gemma 4 31B it 4 bit GPTQ Quantization (W4A16) This is a 4 bit GPTQ quantization of Google's Gemma 4 31B it instruction tuned multimodal model, optimized for deployment on NVIDIA GPUs with vLLM. Model Details Property Value Base Model google/gemma 4 31B it Quantization Method GPTQ W4A16 (asymmetric) Weight Precision 4 bit int4 Activation Precision FP16/BF16 Group Size 128 Quantization Library llm compressor 0.10.0.1 Format compressed tensors (pack quantized) Architecture Gemma4ForConditionalGeneration Layers 60 decoder layers Hidden Size 5376 Context Window 262K tokens Vision Tower SigLIP (27 layers, NOT quantized) Quantized Components Text decoder + projector (vision tower preserved in BF16) Hardware Requirements Verified Deployment : NVIDIA RTX PRO 6000 (96GB VRAM, Blackwell sm120) Actual VRAM Usage : ~35GB with gpu memory utilization: 0.4 (full 256K context) CUDA Version : cu130 (CUDA 13.0) vLLM Version : 0.18.2+ (vllm/vllm openai:gemma4 cu130 Docker image) Note : The 4 bit GPTQ quantization significantly reduces VRAM requirements compared to BF16 (~62GB). The vision tower remains in BF16, but the quantized text decoder enables deployment on high end GPUs with substantial headro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy