Qwen3.6 27B Omnimerge v4 GGUF GGUF quantizations of ManniX ITA/Qwen3.6 27B Omnimerge v4 — the MLP passthrough variant that defends against the Qwen3.6 think policy fragility we discovered. Source dtype is BF16; this repo provides the standard bartowski quant ladder (F16 → IQ2 XXS) for llama.cpp . Source model: ManniX ITA/Qwen3.6 27B Omnimerge v4 (BF16 weights, model card with full benchmarks and methodology). NOT a quant of clean Qwen/Qwen3.6 27B — these GGUFs contain the v4 merge. MTP companion (2× decode speedup): weight identical GGUFs with the MTP head retained for llama.cpp spec type draft mtp self speculative decoding are at ManniX ITA/Qwen3.6 27B Omnimerge v4 MTP GGUF . Quality is statistically indistinguishable from this repo (HE 137/164 ↔ 137/164, GPQA 155/198 ↔ 154/198); aggregate decode is 2.0 2.3 × faster on a single 24 GB GPU. Use that repo for interactive / single request workloads where latency matters. All quants made using imatrix with calibration data v5, the same calibration set bartowski uses for the Qwen3.6 base release — so quality fingerprints are directly comparable to bartowski's Qwen Qwen3.6 27B GGUF repo. Why this merge exists Same base DARE TIES (Omnimer…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy