mlx community/gemma 4 26B A4B it assistant bf16 This model was converted to MLX format from google/gemma 4 26B A4B it assistant using mlx vlm version 0.4.5 . Refer to the original model card for more details on the model. Use with mlx Single request — draft block size 6 : Batched generation — draft block size 3 , use batch generate : About MLX port of Google's Gemma 4 Multi Token Prediction (MTP) drafter for speculative decoding. A small 4 layer assistant drafts several candidate tokens per round; the full Gemma 4 target verifies them in a single forward pass. Output is byte identical to no drafter at temperature=0 . Recommended draft block size : 6 for single requests, 3 for batched generation. See the drafter docs for architecture, supported pairings, performance numbers, and caveats.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy