Gemma 4 12B it AEON Abliterated — MLX 8 bit (mxfp8, near lossless) The near lossless Apple Silicon build of AEON 7/Gemma 4 12B it AEON Abliterated K4 BF16 . Every language decoder linear is mxfp8 (8 bit, group 32); the tied soft capped head and the vision/audio projectors are kept at bf16 . Built and validated on a MacBook Pro M4 Pro (48 GB) . Target hardware: Apple Silicon (M series), best on 24 GB+ unified memory. Full multimodal (text + image + audio) via mlx vlm . Want the smallest build for a 16 GB Mac? See the compact FP4 sibling: … MLXFP4 . This is the maximum fidelity member of the MLX quant grid: measured top 1 token agreement 0.924 and median KL ≈ 0.002 nats against the BF16 source on the model's own greedy output — below the perceptual/sampling noise floor on the typical token. Near lossless , and the recommended build when quality is paramount. ⚡ Quickstart (Apple Silicon) 0 → running on a fresh Mac (no Python, no tools needed) — uv installs a correct Python + the deps for you: Call it like an OpenAI endpoint ( POST http://localhost:8080/v1/chat/completions ) with the request "model" set to the launched id. (While this repo is private, run hf auth login first — or pass…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy