Version 26.05.01 Calibration STEM and Agentic Languages EN ZH HI AR RU JA KO NL FR ES Model Size 17.23 GB Contact Email Parallelism ⚠️ Multi GPU: Please use data parallelism ( data parallel size N ) to scale across GPUs — each replica runs single GPU. Intra model sharding is not yet supported: tensor parallel size 1 and pipeline parallel size 1 both crash at warmup due to TP/PP assumptions not yet handled in the custom DiffusionGemma sampler. Tested with vLLM 0.23.1rc1.dev24+g51ec5cf08 (commit 51ec5cf08 , built with VLLM USE PRECOMPILED=1 ). Hugging Face GitHub Launch Blog Documentation License : Apache 2.0 Authors : Google DeepMind DiffusionGemma is a generative model built by Google DeepMind. Based on the 26B A4B Mixture of Experts (MoE) Gemma 4 architecture, DiffusionGemma generates tokens using discrete diffusion. This open weights model is multimodal, handling text, image, and video inputs to generate text output. Built on a MoE foundation, DiffusionGemma is designed to improve generation speed (tokens per second) while remaining deployable across various hardware environments. DiffusionGemma builds upon the architectural and capability advancements of Gemma 4, introducing sev…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy