mlx community/gemma 4 26B A4B it OptiQ 4bit Built with mlx optiq , the MLX native toolkit to quantize, fine tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptiQ quants · Docs A 4 bit mixed precision MLX quant produced by mlx optiq, the sensitivity aware quantization toolkit for Apple Silicon. Beats stock uniform 4 bit on every benchmark in the six metric Capability Score. A 4 bit mixed precision MLX quant of google/gemma 4 26B A4B it. Per layer bit widths come from a KL divergence sensitivity pass on a six domain calibration mix (prose · reasoning · code · agent · tool call · constraint bearing instructions). Sensitive layers go to 8 bit; robust ones stay at 4 bit. The on disk size is within ~5 % of a stock uniform 4 bit MLX quant. Quantization details Property Value Predominant precision 4 bit Layers at 8 bit (sensitive) 246 Layers at 4 bit (robust) 79 Total quantized layers 325 Group size 64 Calibration mix six domain mix (40 samples × 6 domains) Reference for sensitivity bf16 (auto resolved; falls back to uniform 4 bit if bf16 doesn't fit) Speculative drafter served with mlx community/gemma 4 26B A4B it assistant bf16 via optiq serve dr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy