gemma 3 27b it FP8 Dynamic Model Overview Model Architecture: gemma 3 27b it Input: Vision Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: 2/24/2025 Version: 1.0 Model Developers: Neural Magic Quantized version of google/gemma 3 27b it. Model Optimizations This model was obtained by quantizing the weights of google/gemma 3 27b it to FP8 data type, ready for inference with vLLM = 0.5.2. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM also supports OpenAI compatible serving. See the documentation for more details. Creation This model was created with llm compressor by running the code snippet below as part a multimodal announcement blog. Model Creation Code Evaluation The model was evaluated using lm evaluation harness for OpenLLM v1 text benchmark. The evaluations were conducted using the following commands: Evaluation Commands OpenLLM v1 Accuracy Category Metric google/gemma 3 27b it RedHatAI/gemma 3 27b it FP8 Dynamic Recovery (%) OpenLLM V1 ARC Challenge 72.53% 72.70% 100.24% GSM8K 92.12% 91.51% 99.34% Hellaswag 85.78% 85.69% 99.90% MMLU 77.53% 77…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy