Qwen3.6 27B MTP 4bit This repository contains Multi Token Prediction (MTP) drafter weights split from Qwen/Qwen3.6 27B for use with mlx vlm speculative decoding. This is not a standalone chat or text generation model. Load it as the draft model alongside a compatible Qwen3.6 27B target checkpoint. Use with mlx vlm For local weights: Model Details Model type: qwen3 5 mtp MTP block size: 2 Target architecture: Qwen3.6 27B Precision: MLX affine 4 bit, group size 64 Runtime: MLX / mlx vlm Format: Safetensors with MLX compatible config and tokenizer files The stored tensors use MLX affine 4 bit quantization as described in config.json . Intended Use Use this repo only as a speculative decoding drafter for compatible Qwen3.6 27B checkpoints. The target model verifies drafted tokens, while this MTP model proposes candidate tokens per decoding step. Limitations This checkpoint requires runtime support for Qwen/DeepSeek MTP draft models in mlx vlm . Standard standalone generation through generic Transformers APIs is not expected to work with this repository by itself. Please refer to the upstream Qwen/Qwen3.6 27B model card and license terms for model usage constraints.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy