Gemma 4 12B IT W4A16 AutoRound Quantized Variants This repository hosts 4 bit weight, 16 bit activation (W4A16) quantized variants of google/gemma 4 12B it . The models were quantized using Intel's AutoRound framework tailored specifically for the architectural requirements of the Gemma 4 family. Available formats in the series: Vishva007/gemma 4 12B it W4A16 AutoRound (Standard AutoRound format) Vishva007/gemma 4 12B it W4A16 AutoRound AWQ (AWQ format conversion) Vishva007/gemma 4 12B it W4A16 AutoRound GPTQ (GPTQ format conversion) Architectural Advantage: Gemma 4 12B Unified The Gemma 4 12B Unified model features a ground breaking encoder free multimodal architecture . Unlike traditional vision language models that rely on separate heavy visual/audio encoders (like ViT or Whisper), the 12B Unified model projects raw image patches and audio waveforms directly into the main LLM's embedding space via lightweight linear layers. Because text, image, and audio flow natively into a single decoder only transformer, this model benefits dramatically from weight only quantization, offering minimal multimodal latency and a highly streamlined memory footprint. Quantization Recipe & Environme…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy