Gemma 4 12B it NVFP4 GGUF NVFP4 (Blackwell FP4) quantization of Google's Gemma 4 12B It, a multimodal language model with native vision understanding. This repository contains two files: gemma 4 12b it nvfp4.gguf — Text backbone (48 transformer layers, 3840 hidden dim, 262k context) quantized to NVFP4 mmproj gemma 4 12b it f16.gguf — SigLIP vision encoder + projector at F16 precision (required for image input) About NVFP4 NVFP4 is NVIDIA's native 4 bit floating point format (E4M3 — 1 sign, 4 exponent, 3 mantissa bits) purpose built for Blackwell GPU architecture (RTX 50 series). Unlike block quantized INT4 (Q4 K M) or microscaling MXFP4, NVFP4 operates directly on Blackwell's native tensor core data type, eliminating the dequantization step entirely. Feature NVFP4 Q4 K M MXFP4 Numeric format E4M3 native FP4 INT4 block quantization E2M1 microscaling Block size 32 elements 32 elements 32 elements Effective BPW 4.68 ~4.50 4.72 Dequantization overhead None (native tensor cores) Required on every load Required on every load Hardware acceleration Blackwell (RTX 5060 Ti, 5070, 5090) CUDA cores / CPU CUDA cores / CPU / AMD Dynamic range (max normal) 448 (E4M3) 7 (INT4, symmetric) 30 (E2M1)…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy