Gemma 4 E4B Assistant — GGUF (Atomic Chat) GGUF builds of google/gemma 4 E4B it assistant — the official Gemma 4 Multi Token Prediction (MTP) drafter for google/gemma 4 E4B it . Use it as a speculative decoding draft model alongside the matching Gemma 4 target to get a meaningful decoding speedup at zero quality loss. Approximate size: 78.8M (assistant) / 8B target . [!IMPORTANT] These GGUFs use the custom gemma4 assistant architecture and will not load in stock llama.cpp . They require the atomic llama cpp turboquant fork, which adds: the gemma4 assistant MTP drafter arch (incl. the centroid LM head for E2B/E4B), TurboQuant KV cache quantization ( ctk turbo3 ctv turbo3 ), the mtp head / spec type mtp runtime flags. Loading these files in upstream ggml org/llama.cpp will fail with an unknown architecture error. Files File Quant Size Notes : gemma 4 E4B it assistant.F16.gguf F16 165.8 MB reference (smallest quality loss vs source) gemma 4 E4B it assistant.Q8 0.gguf Q8 0 95.6 MB near lossless 8 bit gemma 4 E4B it assistant.Q5 K M.gguf Q5 K M 76.2 MB balanced k quant gemma 4 E4B it assistant.Q4 K M.gguf Q4 K M 74.9 MB recommended default for speculative decoding draft gemma 4 E4B it a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy