Read our How to Run Gemma 4 QAT Guide! See Unsloth Dynamic 2.0 GGUFs for our quantization benchmarks. Jun 9 Update: Added MTP support. See our MTP Guide . Gemma 4 can now be run and fine tuned in Unsloth Studio . Read our guide . See all versions of Gemma 4 QAT (GGUF, 16 bit etc.) in our collection . Example of Gemma 4 E4B (4 bit GGUF) running in Unsloth Studio with tool calling: Run with MTP (speculative decoding) This model ships a Multi Token Prediction drafter at the repo root ( mtp gemma 4 12B it.gguf , a near lossless smart Q4 0). A recent llama.cpp auto discovers it from hf , so you do not pass model draft : The drafter shares the target's KV cache and does not change the output (the target verifies every drafted token). See the MTP/ folder for the other precisions and explicit usage. Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind [!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantiz…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy