Gemma 4 26B A4B Assistant GGUF GGUF quantizations converted from google/gemma 4 26B A4B it qat q4 0 unquantized assistant . Tested with llama.cpp b9549 (Gemma 4 MTP support). Update Added experimental IQ quantizations with Q4 embeddings (token embd.weight = Q4 0). Recommendations Q4 0 q4emb — recommended for most users Q8 0 — for users with spare VRAM Files gemma 4 26B A4B it assistant f16.gguf gemma 4 26b A4B it assistant Q4 0.gguf gemma 4 26b A4B it assistant Q4 0 q4emb.gguf (closest to pure Q4 QAT layout) gemma 4 26b A4B it assistant IQ4 NL q4emb.gguf gemma 4 26b A4B it assistant IQ3 M q4emb.gguf (smallest that still works) gemma 4 26b A4B it assistant Q8 0.gguf Q4 Embedding Variant Q4 0 q4emb is an experimental quantization where token embd.weight is kept in Q4 0 instead of Q6 K precision quantization typically used by llama.cpp. This follows a similar approach to recent QAT experiments for Gemma models, where preserving the original Q4 trained embedding format may better match the intended QAT behavior. Initial testing showed similar draft acceptance rates to the default Q4 0 quant, with a small speed advantage, though more benchmarking is needed. Example Recommended values: s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy