Gemmable 4 12B Gemmable 4 12B is a GGUF export of Gemma 4 12B fine tuned on Fable 5 style reasoning and assistant traces. Highlights Base model: google/gemma 4 12B Format: GGUF Training style: Fable 5 style reasoning and assistant traces Distribution: fp16 GGUF plus matching assistant GGUFs for each quant Intended use: local inference, coding, reasoning, and assistant workflows How to use llama.cpp Standard load: Speculative / draft MTP load: Use the matching fp16 or quantized main file with its mtp companion. LM Studio 1. Search this repo, download target + mtp file. 2. Load target. 3. Load settings → Speculative Decoding → select mtp file file. (Requires LM Studio with am17an's PR merged or custom llama.cpp runtime. As of 2026 05, mainline LM Studio runtime doesn't yet have draft mtp for Gemma 4 — track upstream merge.) GGUF / local inference notes gemmable 4 12b fp16.gguf is the standard fp16 main model. gemmable 4 12b fp16 mtp.gguf is the matching fp16 assistant / draft file. Quantized pairs follow the same pattern, for example gemmable 4 12b Q4 K M.gguf and gemmable 4 12b Q4 K M mtp.gguf . Keep the paired files in the same Hugging Face repository if you upload them. Limitation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy