Llamacpp imatrix Quantizations of gemma 3 12b it by google Using llama.cpp release b4877 for quantization. Original model: https://huggingface.co/google/gemma 3 12b it All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Vision This model has vision capabilities, more details here: https://github.com/ggml org/llama.cpp/pull/12344 After building with Gemma 3 clip support, run the following command: Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description mmproj gemma 3 12b it f32.gguf f32 1.69GB false F32 format MMPROJ file, required for vision. mmproj gemma 3 12b it f16.gguf f16 854MB false F16 format MMPROJ file, required for vision. gemma 3 12b it bf16.gguf bf16 23.54GB false Full BF16 weights. gemma 3 12b it Q8 0.gguf Q8 0 12.51GB false Extremely high quality, generally unneeded but max available quant. gemma 3 12b it Q6 K L.gguf Q6 K L 9.90GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . gemma 3 12b it Q6 K.gguf Q6 K 9.66GB false Very high quality, near perfect, recommended . ge…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy