gemma 3 4b it GPTQ 4b 128g Model Overview This model was obtained by quantizing the weights of gemma 3 4b it to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. Only the weights of the linear operators within language model transformers blocks are quantized. Vision model and multimodal projection are kept in original precision. Weights are quantized using a symmetric per group scheme, with group size 128. The GPTQ algorithm is applied for quantization. Model checkpoint is saved in compressed tensors format. Evaluation This model was evaluated on the OpenLLM v1 benchmarks. Model outputs were generated with the vLLM engine. Model ArcC GSM8k Hellaswag MMLU TruthfulQA mc2 Winogrande Average Recovery : : : : : : : : : : : : : : : : gemma 3 4b it 0.6084 0.7528 0.7497 0.5832 0.5189 0.7072 0.6534 1.0000 gemma 3 4b it INT4 (this) 0.5879 0.7210 0.7358 0.5650 0.4863 0.6811 0.6295 0.9635 Reproduction The results were obtained using the following commands: Usage To use the model in transformers update the package to stable release of Gemma3: pip install git+https://github.com/hugging…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy