Llamacpp imatrix Quantizations of gemma 3 4b it by google Using llama.cpp release b4877 for quantization. Original model: https://huggingface.co/google/gemma 3 4b it All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Vision This model has vision capabilities, more details here: https://github.com/ggml org/llama.cpp/pull/12344 After building with Gemma 3 clip support, run the following command: Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description mmproj gemma 3 4b it f32.gguf f32 1.68GB false F32 format MMPROJ file, required for vision. mmproj gemma 3 4b it f16.gguf f16 851MB false F16 format MMPROJ file, required for vision. gemma 3 4b it bf16.gguf bf16 7.77GB false Full BF16 weights. gemma 3 4b it Q8 0.gguf Q8 0 4.13GB false Extremely high quality, generally unneeded but max available quant. gemma 3 4b it Q6 K L.gguf Q6 K L 3.35GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . gemma 3 4b it Q6 K.gguf Q6 K 3.19GB false Very high quality, near perfect, recommended . gemma 3 4b i…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy