This is a NVFP4 quantized variant of google/gemma 4 E4B it . Weights and activations of the core transformer linear layers have been quantized to NVFP4 (W4A4) using NVIDIA Model Optimizer with mixed precision AutoQuantize. Per Layer Embeddings (PLE), the vision tower, audio tower, and multimodal projector are kept in BF16 to preserve multimodal quality. Requirements Full W4A4 inference requires an NVIDIA Blackwell GPU (RTX 50 series / RTX PRO 6000 Blackwell / B series). vLLM version must be higher than or equal to v0.19.0. Cuda version must be higher than or equal to 12.9. ⚠️ On pre Blackwell GPUs (Ada, Hopper, Ampere) the model will load but FP4 activation kernels won't fire, losing the W4A4 throughput benefit. Usage with vLLM Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on small models) and generating text output. This release includes open weights models in both pre trained and instruction tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual suppor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy