Base model: google/gemma 4 E4B it Gemma 4 E4B , self quantized to GGUF by Atomic Chat. Built straight from Google's original weights with a per tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. Highlights 4.5B effective (8B with embeddings) parameters : the weights this repo quantizes. Context length : 128K tokens, as published by Google. 42 layers : Dense decoder, hybrid sliding window (512) and global attention. Modalities : Text, Image, Audio. Full imatrix ladder : every quant is calibrated with an importance matrix, published here alongside the quants. Reasoning : All models in the family are designed as highly capable reasoners, with configurable thinking modes. Diverse & Efficient Architectures : Offers Dense and Mixture of Experts (MoE) variants of different sizes for scalable deployment. [!NOTE] These GGUFs are self quantized from the original weights , not a repack. The importance matrix keeps low bit quants closer to the full precision model. [!IMPORTANT] Always pass jinja so the Gemma 4 E4B chat template is applied. Without it the model can emit malformed turns. Model Overview Property Value Base model google/gemma 4 E4B it P…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy