Base model: google/gemma 4 31B it assistant Gemma 4 31B It Assistant , self quantized to GGUF by Atomic Chat. Built straight from Google's original weights with a per tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. Highlights 30.7B parameters : the weights this repo quantizes. Context length : 256K tokens, as published by Google. 60 layers : Dense decoder, hybrid sliding window (1024) and global attention. Modalities : the base model handles Text, Image; this repo ships text only quants, it carries no vision projector. Full imatrix ladder : every quant is calibrated with an importance matrix. Reasoning : All models in the family are designed as highly capable reasoners, with configurable thinking modes. Diverse & Efficient Architectures : Offers Dense and Mixture of Experts (MoE) variants of different sizes for scalable deployment. [!NOTE] These GGUFs are self quantized from the original weights , not a repack. The importance matrix keeps low bit quants closer to the full precision model. [!IMPORTANT] Always pass jinja so the Gemma 4 31B It Assistant chat template is applied. Without it the model can emit malformed turns. Model Overvi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy