mlx community/Qwen3.5 9B OptiQ 4bit Built with mlx optiq , the MLX native toolkit to quantize, fine tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptIQ quants · Docs A 4 bit mixed precision MLX quant produced by mlx optiq, the sensitivity aware quantization toolkit for Apple Silicon. Beats stock uniform 4 bit on every benchmark in the six metric Capability Score. A 4 bit mixed precision MLX quant of Qwen/Qwen3.5 9B. Per layer bit widths come from a KL divergence sensitivity pass on a six domain calibration mix (prose · reasoning · code · agent · tool call · constraint bearing instructions). Sensitive layers go to 8 bit; robust ones stay at 4 bit. The on disk size is within ~5 % of a stock uniform 4 bit MLX quant. Quantization details Property Value Predominant precision 4 bit Layers at 8 bit (sensitive) 132 Layers at 4 bit (robust) 116 Total quantized layers 248 Group size 64 Calibration mix six domain mix (40 samples × 6 domains) Reference for sensitivity bf16 (auto resolved; falls back to uniform 4 bit if bf16 doesn't fit) Bundled MTP head mtp.safetensors (4 bit projections, BF16 norms), enables 1.4× decode via optiq serve mtp We follow…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy