gemma 4 31B it NVFP4 turbo vision A more compact NVFP4 quantization of google/gemma 4 31B it that keeps the vision tower intact so multimodal (image + text) input still works. This is a vision preserving variant of LilaRest's excellent gemma 4 31B it NVFP4 turbo . Same RTN on attention recipe; the only difference is that the BF16 vision tower, vision embedding projection, and vision config are retained, and the architecture stays Gemma4ForConditionalGeneration . For the recipe rationale, hardware requirements, and architecture compatibility matrix, see LilaRest's model card — it's the canonical reference. What's quantized vs. preserved Component Treatment Source language model.layers. .mlp.{up,gate,down} proj NVFP4 (calibrated PTQ) Inherited from nvidia/Gemma 4 31B IT NVFP4 language model.layers. .self attn.{q,k,v,o} proj NVFP4 (RTN, no calibration) New in this variant language model.embed tokens (tied to lm head) BF16 (untouched) — layernorm / norm BF16 (untouched) — vision tower. BF16 (untouched) Retained from NVIDIA source embed vision. BF16 (untouched) Retained from NVIDIA source The RTN step is byte deterministic from the source weights — no calibration dataset, no forward pas…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy