Gemma 4 12B Instruction Tuned — GGUF (multimodal) Community GGUF mirror of google/gemma 4 12B it for local, encoder free multimodal AI on consumer hardware (~16 GB VRAM). Announced June 2026: Google blog · Developer guide Parameters ~12B dense Modalities Text, vision , audio (native in backbone) License Apache 2.0 Architecture Encoder free (no separate vision/audio towers) Context See upstream config Vision in GGUF Requires mmproj .gguf alongside main weights Why this repo exists One download hub for all major quants (K quants, IQ, Q8, mmproj). Fast Hub side sync from bartowski/gemma 4 12B it GGUF — no re upload from your laptop. Documented use cases for contributors: gemma 4 12b local (agents, LiteRT, llama.cpp, MLX). Apple Silicon MLX quants: Edmon02/gemma 4 12B it MLX Available files See gguf manifest.json for the live file list. Essential tier (recommended) File Use gemma 4 12B it Q4 K M.gguf Best balance — 16 GB laptops gemma 4 12B it Q5 K M.gguf Higher quality gemma 4 12B it Q6 K.gguf / Q8 0 Max quality gemma 4 12B it Q3 K M.gguf Tighter VRAM gemma 4 12B it Q2 K.gguf Minimum size gemma 4 12B it IQ4 XS.gguf / IQ4 NL.gguf IQ variants mmproj gemma 4 12B it f16.gguf Required for…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy