Gemma 4 12B IT Assistant GGUF Q8 0 This repository contains a GGUF Q8 0 conversion of google/gemma 4 12B it assistant . This GGUF is intended to be used as an MTP / speculative draft model with a compatible Gemma 4 main model. At the time of upload, this requires the gemma4 mtp llama.cpp branch: https://github.com/am17an/llama.cpp/tree/gemma4 mtp Usage This model is not intended to be run as a standalone chat model. It is intended to be loaded as the draft model with md . Required llama.cpp branch: Example llama server command: Benchmark Benchmark was run with llama server.exe from the gemma4 mtp llama.cpp branch. Setting Value Main model gemma 4 12B it Q6 K.gguf Draft model gemma 4 12B it assistant Q8 0.gguf Required branch am17an/llama.cpp gemma4 mtp Runs 3 prompts x 5 measured repeats per mode, with 1 warmup per prompt Generation length 256 tokens Context 262144 GPU layers 99 Flash attention on KV cache Q8 0 / Q8 0 Temperature 0 Generation throughput Mode Short gen tok/s Medium gen tok/s Long gen tok/s Mean gen tok/s Mean speedup vs baseline : : : : : Baseline 52.08 51.94 51.10 51.71 MTP draft n=1 57.79 56.27 58.21 57.42 +11.0% MTP draft n=2 57.83 59.45 61.54 59.61 +15.3% MTP dr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy