Gemma 4 E2B Instruct — W4A16 Quantized AutoRound This repository hosts W4A16 INT4 quantized versions of google/gemma 4 E2B it , a multimodal mixture of experts model supporting text, vision, and audio inputs. Two quantized variants are available: Note on MTP / Speculative Decoding: If you want to use the official speculative decoding assistant model ( google/gemma 4 E2B it assistant ) for MTP support, it is recommended to use the vllm/vllm openai:gemma4 0505 cu129 Docker image, which includes newer Gemma 4 support and decoding patches. Due to INT4 quantization, the assistant acceptance rate may be lower compared to the original unquantized google/gemma 4 E2B it model. Variant Method Repo AutoRound (RTN) intel/auto round Vishva007/gemma 4 E2B it W4A16 AutoRound GPTQ AutoGPTQ Vishva007/gemma 4 E2B it W4A16 AutoRound GPTQ Note: Only the language model (LM) layers are quantized to INT4. The vision tower, audio tower, and multimodal projectors are kept at full precision (BF16) to preserve multimodal quality. Quantization Details Parameter Value Base model google/gemma 4 E2B it Quantization scheme W4A16 (INT4 weights, BF16 activations) Group size 128 Symmetric Yes Calibration samples 256…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy