Gemma 4 12B IT QAT Assistant MTP Q8 0 GGUF This repository contains a GGUF conversion of the official Google Gemma 4 12B IT QAT assistant/drafter checkpoint. Source checkpoint: google/gemma 4 12B it qat q4 0 unquantized assistant Output file: gemma 4 12B it qat assistant MTP Q8 0.gguf Quantization: Q8 0 Format: GGUF Intended runtime: llama.cpp with Gemma 4 MTP / draft model support This is not a standalone chat model. It is an assistant / drafter / MTP head intended to be used together with a matching Gemma 4 12B IT QAT target model for speculative decoding. File File Description gemma 4 12B it qat assistant MTP Q8 0.gguf Q8 0 GGUF conversion of the Gemma 4 12B QAT assistant checkpoint Source Converted from the official Google checkpoint: google/gemma 4 12B it qat q4 0 unquantized assistant Usage This GGUF is a draft / assistant / MTP model, not a standalone chat model. It must be loaded together with a matching Gemma 4 12B IT QAT target model. llama server example: llama server \ m gemma 4 12B it qat UD Q4 K XL.gguf \ model draft gemma 4 12B it qat assistant MTP Q8 0.gguf \ spec type draft mtp \ spec draft n max 4 Conversion Converted with llama.cpp using Gemma 4 assistant / MTP s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy