RedHatAI/diffusiongemma 26B A4B it FP8 dynamic This model is an FP8 quantized version of google/diffusiongemma 26B A4B it. The model has both weights and activations quantized to FP8 using vllm/llm compressor and in the compressed tensors format. It was evaluated on several tasks to assess its quality in comparison to the unquantized model using vLLM. Deployment Creation Accuracy The following metrics were generated when serving the quantized model with vLLM on a single B200 GPU. Benchmark google/diffusiongemma 26B A4B it RedHatAI/diffusiongemma 26B A4B it FP8 dynamic Recovery (%) : : : AIME 2025 0.437 0.423 96.8% GPQA Diamond 0.641 0.657 102.5% IFEval 0.879 0.862 98.1% GSM8K 0.943 0.942 99.9% MMLU 0 Shot 0.539 0.505 93.7% Thinking AIME 2025 0.650 0.660 101.5% GPQA Diamond 0.698 0.689 98.7% GSM8K 0.951 0.952 100.1%
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy