Qwen3.6 27B Omnimerge v4 — IQ4 XS (Mixed Bit, 12.6 GiB) Runtime note: This quant was built and tested with ikawrakow's ik llama.cpp fork. It has not been tested on the mainline ggml/llama.cpp. For best results and full feature support (mixed bit quant loading), use ik llama.cpp . Base model: OmniMerge v4 on Qwen3.6 27B Size: 12.76 GiB (13.7 GB, 13,054 MiB on disk) — 4.06 bpw VRAM: Fits 16 GB GPUs comfortably with 32k+ context License: Apache 2.0 A custom importance matrix guided mixed bit quantization of the OmniMerge v4 merge. Achieves lower perplexity than the official uniform IQ4 XS release while being 1.3 GiB smaller — a clear win for targeted bit allocation at this size tier. Results Perplexity (wiki.test.raw, 580 chunks, n ctx=512) Model PPL Size Δ from official IQ4 XS IQ4 XS (mixed bit, ours) 6.864 ± 0.045 12.76 GiB −0.055 IQ4 XS (uniform, official) 6.919 ± 0.045 14.05 GiB — F16 (estimated) ~6.70 50.11 GiB — Both quants lose ~0.15–0.20 PPL from the F16 baseline. The mixed bit recipe recovers ~25% of that loss compared to uniform IQ4 XS, while using 9% less disk space KLD vs Upstream IQ4 XS Metric Value Interpretation Mean KLD 0.0371 ± 0.0015 Low — distributions are very simi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy