Hugging Face GitHub Launch Blog Documentation Technical Report License : Apache 2.0 Authors : Google DeepMind [!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q4 0): Half precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q4 0): Ready to deploy formats for broad ecosystem compatibility. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B. Mobile optimized (wNa8o8): A custom schema engineered explicitly for mobile hardware efficiency. It features targeted 2 bit decoding layers, optimized KV caches, and static activations to maximize VRAM savings. Available for Gemma 4 E2B and E4B. Compressed Tensors (w4a16): QAT checkpoints serialized in the compressed tensors format for native, optimized inference with vLLM. Available for Gemma 4 E2B, E4B, 12B, and 31B. Gemma is a family…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy