mlx community/gemma 4 12B it OptiQ 4bit Built with mlx optiq , the MLX native toolkit to quantize, fine tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. Try the Lab · All OptiQ quants · Docs A 4 bit mixed precision MLX quant produced by mlx optiq, the sensitivity aware quantization toolkit for Apple Silicon. Scores +6.40 over stock uniform 4 bit on the six metric Capability Score, the second largest mixed precision gain in the Gemma 4 lineup. A 4 bit mixed precision MLX quant of google/gemma 4 12B it, the unified (text+vision+audio) Gemma 4. This artifact is the text inference path : the language tower is quantized and the vision/audio towers are dropped during conversion. Per layer bit widths come from a KL divergence sensitivity pass on a six domain calibration mix (prose, reasoning, code, agent, tool call, constraint bearing instructions). Sensitive layers go to 8 bit; robust ones stay at 4 bit. Quantization details Property Value Predominant precision 4 bit Layers at 8 bit (sensitive) 156 Layers at 4 bit (robust) 172 Total quantized layers 328 Average bits per weight 5.22 Group size 64 Calibration mix six domain mix (40 samples × 6 domains) Reference for…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy