Gemma 3 4B Instruction tuned QAT compressed tensors This checkpoint was converted from https://huggingface.co/google/gemma 3 4b it qat q4 0 gguf to compressed tensors format and BF16 dtype (hence, not lossess). You can run this with vLLM Below is the original model card. Gemma 3 model card Model Page : Gemma [!Note] This repository corresponds to the 4B instruction tuned version of the Gemma 3 model in GGUF format using Quantization Aware Training (QAT). The GGUF corresponds to Q4 0 quantization. Thanks to QAT, the model is able to preserve similar quality as bfloat16 while significantly reducing the memory requirements to load the model. You can find the half precision version here. Resources and Technical Documentation : [Gemma 3 Technical Report][g3 tech report] [Responsible Generative AI Toolkit][rai toolkit] [Gemma on Kaggle][kaggle gemma] [Gemma on Vertex Model Garden][vertex mg gemma3] Terms of Use : [Terms][terms] Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state of the art open models from Google, built from the same research and technology used to create the Gemin…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy