gemma 4 31B it FP8 gemma 4 31B it FP8 is an FP8 compressed evolution of google/gemma 4 31B it. This variant leverages BF16 · F8 E4M3 precision formats to significantly reduce memory footprint and improve inference efficiency while maintaining strong output quality. gemma 4 31B it from Google is the flagship dense model in the Gemma 4 family, featuring 31 billion parameters optimized for workstation and server deployment with a massive 256K context window, supporting text and images (variable aspect ratios and resolutions) along with advanced agentic capabilities such as step by step thinking modes, multilingual OCR and handwriting recognition, document and PDF parsing, UI and screen analysis, chart comprehension, and precise object detection with pointing. Designed to bridge edge and cloud performance, the instruction tuned variant delivers frontier level reasoning rivaling proprietary models 5 to 10× larger across coding, math, multilingual tasks (140+ languages), and multimodal workflows while maintaining Google's production grade safety alignments for enterprise use. With Apache 2.0 licensing and optimizations for NVIDIA and AMD GPUs via vLLM and llama.cpp, it enables high quali…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy