RedHatAI/diffusiongemma 26B A4B it NVFP4 This model is an NVFP4 quantized version of google/diffusiongemma 26B A4B it. The model has both weights and activations quantized to NVFP4 using vllm/llm compressor and in the compressed tensors format. It was evaluated on several tasks to assess its quality in comparison to the unquantized model using vLLM. Deployment Creation Accuracy The following metrics were generated when serving the quantized model with vLLM on a single B200 GPU. Benchmark google/diffusiongemma 26B A4B it RedHatAI/diffusiongemma 26B A4B it NVFP4 Recovery (%) : : : AIME 2025 0.437 0.427 97.7% GPQA Diamond 0.641 0.644 100.5% IFEval 0.879 0.866 98.5% GSM8K 0.943 0.943 100.0% MMLU 0 Shot 0.539 0.616 114.3% Thinking AIME 2025 0.650 0.637 98.0% GPQA Diamond 0.698 0.677 97.0% GSM8K 0.951 0.952 100.1%
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy