๐ MoQ: Mixture of Quants This is the GGUF version of gemma 4 12b it qat model . It uses MoQ to improve performance over uniform quantization. MoQ (Mixture of Quants) is a smart way to shrink AI models without losing their "brainpower." Unlike old methods that treat every part of the model the same, MoQ identifies the most important parts and keeps them high quality, while heavily compressing the rest to save space. The result? A model that punches significantly above its weight class. Benjamin Marie evaluated MoQ GGUFs ("Mixture of Quants") against Unsloth Dynamic (UD) quants, focusing on low bit versions below 4 bits on average โ the range where GGUF models typically struggle most. Results: At similar bits per weight (Bpw), MoQ outperforms Unsloth Dynamic quants by ~10% on benchmarks, while also being roughly 2ร more token efficient on average. "MoQ models are much better than UD quants on benchmarks, and they are also more token efficient." Folder Link BPW Total Size Description : : : : : : ๐ MoQ Quants 3.0 4.06 GB ๐ MoQ Quants 3.25 4.81 GB ๐ MoQ Quants 3.5 5.05 GB ๐ MoQ Quants 3.75 5.63 GB ๐ MoQ Quants 4.0 6.38 GB ๐ MoQ Quants 4.25 6.46 GB ๐ MoQ Quants 4.5 6.54 GB ๐ MoQโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy