This is a preliminary version (and subject to change) of the FP8 quantized google/gemma 4 31B it model, distributed by BC Card . This work was carried out with reference to the Red Hat AI approach and validation methodology. The model has both weights and activations quantized to FP8 with vllm project/llm compressor . This model requires a nightly vllm wheel. For the reference installation and execution flow, see the Red Hat AI / vLLM based guidance: https://docs.vllm.ai/projects/recipes/en/latest/Google/Gemma4.html installing vllm On a single B200: This is a preliminary version (and subject to change) of the FP8 quantized google/gemma 4 31B it model. The model has both weights and activations quantized to FP8 with vllm project/llm compressor. This model requires a nightly vllm wheel, see install instructions at https://docs.vllm.ai/projects/recipes/en/latest/Google/Gemma4.html installing vllm On a single B200: Original: FP8:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy