NVFP4 Quantized RedHatAI/gemma 4 12B it NVFP4 This is a preliminary version (and subject to change) of FP8 Dynamic quantized google/gemma 4 12B it model. The model has both weights and activations quantized to FP8 Dynamic format with vllm project/llm compressor. It is compatible and tested against vllm nightly. Creation Script Run this script with this LLM Compressor PR and latest transformers to quantize the model using iMatrix quantization Preliminary Evaluations 1) GSM8K Platinum 2) Wikitext PPL Evals:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy