Gemma 4 31B it NVFP4A16 Quantization (Weight only FP4) This is a NVFP4A16 quantization of Google's Gemma 4 31B it instruction tuned multimodal model, optimized for deployment on NVIDIA Blackwell GPUs with vLLM. Model Details Property Value Base Model google/gemma 4 31B it Quantization Method NVFP4A16 (weight only FP4) Weight Precision FP4 (4 bit floating point) Activation Precision BF16 (weight only quantization) Group Size 16 Quantization Library llm compressor 0.10.0.1 Format compressed tensors (nvfp4 pack quantized) Architecture Gemma4ForConditionalGeneration Layers 60 decoder layers Hidden Size 5376 Context Window 256K tokens Vision Tower SigLIP (27 layers, preserved in BF16) Quantized Components Text decoder + projector (vision tower preserved in BF16) Hardware Requirements Verified Deployment : NVIDIA RTX PRO 6000 (96GB VRAM, Blackwell sm120) Actual VRAM Usage : ~35GB with gpu memory utilization: 0.4 (full 256K context) CUDA Version : cu130 (CUDA 13.0) vLLM Version : 0.18.2+ (tested on vllm/vllm openai:gemma4 cu130 Docker image) Note : NVFP4A16 is a weight only quantization format that preserves activations in BF16. The vision tower remains in BF16, but the quantized text dec…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy