Read our How to Run NVFP4 quants Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. Gemma 4 12B can now be run and fine tuned in Unsloth Studio . Read our guide . See all versions of Gemma 4 (GGUF, 16 bit etc.) in our collection . Example of Gemma 4 E4B (4 bit GGUF) running in Unsloth Studio with tool calling: Run in vLLM Do not use the Marlin backend (around 2x slower); let vLLM auto select the NVFP4 kernel. Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind [!Note] This model card is for the Gemma 4 12B Unified model, which is part of the Gemma 4 family of open models. Built with the same multimodal functionality as Gemma 4 E2B and E4B (text, audio, image, and video inputs), it brings native audio and vision understanding directly to local environments without the need for separate encoders. This unified approach to multimodality makes the model encoder free, offering a deployment size that is perfect for consumer devices and streamlined local execution. Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy