Gemma 3 Quantized Models This repository contains W4A16 quantized versions of Google's Gemma 3 instruction tuned models, making them more accessible for deployment on consumer hardware while maintaining good performance. Models abhishekchohan/gemma 3 27b it quantized W4A16 abhishekchohan/gemma 3 12b it quantized W4A16 abhishekchohan/gemma 3 4b it quantized W4A16 Repository Structure Quantization Details These models use W4A16 quantization via LLM Compressor: Weights quantized to 4 bit precision Activations use 16 bit precision Significantly reduced memory requirements Usage with vLLM License These models are subject to the Gemma license. Users must acknowledge and accept the license terms before using the models. Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy