gemma 4 26B A4B it MTP GGUF Speculative decoding bundle for google/gemma 4 26B A4B it using Google's official MTP drafter ( gemma 4 26B A4B it assistant ), packaged for llama.cpp and LM Studio . Achieves ~ 1.55x throughput vs no MTP on Apple M5 Max (Q2 K drafter, n=3, thinking off). Files File Role Size gemma 4 26B A4B it Q8 0.gguf target (26B A4B MoE, 128 experts) 26 GB gemma 4 26B A4B it assistant Q2 K.gguf drafter (recommended) 278 MB gemma 4 26B A4B it assistant Q4 K M.gguf drafter (balanced) 310 MB gemma 4 26B A4B it assistant Q8 0.gguf drafter (high precision) 440 MB gemma 4 26B A4B it assistant F16.gguf drafter (reference) 816 MB Target re hosted from unsloth/gemma 4 26B A4B it GGUF (Unsloth Dynamic 2.0 quant, Apache 2.0). Drafter built from google/gemma 4 26B A4B it assistant via llama.cpp PR 23398 (am17an, WIP). Requirements llama.cpp built from am17an's gemma4 mtp branch — Gemma 4 MTP not yet merged to master. Run Or pull directly: LM Studio 1. Search this repo, download target + drafter. 2. Load target. 3. Load settings → Speculative Decoding → select drafter file. (Requires LM Studio with am17an's PR merged or custom llama.cpp runtime. As of 2026 05, mainline LM Studio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy