gemma 4 26B A4B it qat W4A16 Lossless Q4 0 GGUF → compressed tensors pack quantized W4A16 model for vLLM deployment. What This Is Base model: google/gemma 4 26B A4B it Source format: Q4 0 GGUF (QAT trained, exported by llama.cpp) Target format: compressed tensors pack quantized Quantization: 4 bit symmetric integer, group wise (group size=32) Scale dtype: bfloat16 (or float16 via dtype ) Marlin padding: MoE and Dense MLP blocks padded to min thread k=128 (Dense MLP skippable via skip dense mlp ) Vision tower: copied from unquantized ref source (unquantized bfloat16) k eq v handling: v proj absent for full attention layers (5, 11, 17, 23, 29) Conversion The conversion is lossless at the nibble level — every Q4 0 scale and quantized weight value is preserved bit for bit from the GGUF source. Arguments Flag Default Description gguf path (required) Path to Q4 0 GGUF file unquantized ref (required) Unquantized bf16 model (config + tokenizer + non Q4 0 weights) output dir (required) Output directory dtype bfloat16 Dtype for unquantized tensors and scales ( bfloat16 or float16 ) group size 32 CT quantization group size no marlin pad off Skip Marlin K dimension padding skip dense mlp off K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy