Qwen3.6 27B MTPLX Optimized Speed Run this with MTPLX MTPLX is an MLX native runtime for native Multi Token Prediction speculative decoding on Apple Silicon. Up to 2.24× faster decode at real coding temperatures ( temp=0.6 / top p=0.95 / top k=20 ) using the model's own built in MTP heads — no external drafter, no greedy hack. Project: github.com/youssofal/MTPLX Other MTPLX checkpoints: Qwen3.5 4B MTPLX Optimized Speed — small 4 bit speed test Qwen3.5 4B Optimized MTPLX — small 8 bit This is the speed checkpoint for MTPLX v0.1.0 preview. It combines the and embeds mtplx runtime.json so MTPLX can run it as the optimized speed variant. Target repository: Youssofal/Qwen3.6 27B MTPLX Optimized Speed Runtime MTPLX v0.1.0 preview supports Qwen3 Next MTP only. This model is verified for the qwen3 next mtp backend. Sources And Attribution Component Source Revision License : Base model Qwen/Qwen3.6 27B 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9 Apache 2.0 MTPLX conversion and runtime contract youssofal/mtplx v0.1.0 preview.1 Apache 2.0 The trunk is the MLXCommunity flat 4 bit affine conversion of compressed tensors checkpoint at the revision listed above. Verification mtplx runtime.json recor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy