pipeline tag: text generation base model: google/diffusiongemma 26B A4B it license: apache 2.0 license name: apache license 2.0 license link: https://ai.google.dev/gemma/apache 2 tags: nvidia ModelOpt DiffusionGemma 26B A4B IT quantized NVFP4 nvfp4 Model Overview Description: DiffusionGemma 26B A4B IT is an open weights multimodal generative model developed by Google DeepMind that processes text, image, and video inputs to produce text output via discrete diffusion. Built on the Gemma 4 26B A4B Mixture of Experts (MoE) architecture with 25.2B total parameters and 3.8B active parameters, the model employs an encoder decoder design with bidirectional attention that generates tokens in parallel 256 token blocks, enabling high speed generation exceeding 1,100 tokens per second at low batch sizes on NVIDIA Hopper H100 (FP8). DiffusionGemma 26B A4B IT supports a 256K token context window, configurable thinking (reasoning) mode, native function calling, and multilingual inference across 35+ languages. The NVIDIA DiffusionGemma 26B A4B IT NVFP4 model is quantized with Model Optimizer. This model is ready for commercial and non commercial use. Third Party Community Consideration This model…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy