Bonsai 27B — Standard GGUF Ladder (Q3 K M → Q8 0) Imatrix GGUF quants of prism ml's Bonsai 27B (Qwen3.6 27B backbone, 262K context, hybrid attention, vision), covering the quality tiers between the official QAT releases and F16. Which repo should you use? If you have ≤8GB: use prism ml's official QAT quants, not these. Q1 0 (3.8GB) — 89.5% of F16 quality at 1.125 bits Ternary (7.2GB) — 95% of F16 Those are quantization aware trained with custom kernels — at their sizes, they beat anything post training quantization can produce, including anything in this repo. This repo covers the gap above them : the standard ladder for 12–32GB setups where you want maximum quality per GB, quantized from prism ml's own F16 GGUF with an importance matrix (63KB coding/reasoning calibration corpus). Quants File Size Measured (GB10, 273GB/s) + DSpark drafter Q8 0 28.6 GB 7.8 tok/s — Q6 K 22.1 GB 9.1 tok/s — Q5 K M 19.2 GB 10.2 tok/s — Q4 K M 16.6 GB 12.0 tok/s 15.7 tok/s (+31%) IQ4 XS 15.1 GB 13.7 tok/s — Q3 K M 13.3 GB 13.5 tok/s — All tiers individually smoke tested (coherent code generation, chat template engages via jinja ). Q8 0 quantized without imatrix (unneeded at 8 bit); all others use the in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy