gemma 4 12B it qat assistant 4bit This repository contains Multi Token Prediction (MTP) drafter weights split from google/gemma 4 12B it qat q4 0 unquantized assistant for use with mlx vlm speculative decoding. This is not a standalone chat or text generation model. Load it as the draft model alongside a compatible Gemma 4 12B target checkpoint. Use with mlx vlm For local weights: Model Details Model type: gemma4 unified assistant Target architecture: Gemma 4 12B Precision: 4bit Runtime: MLX / mlx vlm Format: Safetensors with MLX compatible config and tokenizer files The stored tensors are 4bit MLX compatible drafter weights. Intended Use Use this repo only as a speculative decoding drafter for compatible Gemma 4 12B checkpoints. The target model verifies drafted tokens, while this MTP model proposes candidate tokens per decoding step. Limitations This checkpoint requires runtime support for Gemma 4 MTP draft models in mlx vlm . Standard standalone generation through generic Transformers APIs is not expected to work with this repository by itself. Please refer to the upstream google/gemma 4 12B it model card and license terms for model usage constraints.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy