⚡ Gemma 4 31B IT NVFP4 Turbo ➡️ Sponsored by Clearbox . Turn Reddit into pipeline. A repackaged nvidia/Gemma 4 31B IT NVFP4 that is 68% smaller in GPU memory and ~2.5× faster than the base model, while retaining nearly identical quality (1 3% loss). Fits on a single RTX 5090 (🎉). It fully leverages NVIDIA Blackwell FP4 tensor cores (RTX 5090, RTX PRO 6000, B200, and other SM 12.0+ GPUs with ≥20 GB VRAM) for ~2× higher concurrent throughput than other quants like prithivMLmods/gemma 4 31B it NVFP4 or cyankiwi/gemma 4 31B it AWQ 4bit. This variant is text only , video/audio weights and encoders have been stripped. If you need video/audio support open an issue or PR. Benchmark [!NOTE] RTX PRO 6000, vllm bench @ 1K input / 200 output tokens. See bench.sh. Note: We also ran ⚡Turbo benchmark on RTX 5090, and it performed exactly the same because at 16k context, the performance is not limited by the GPU memory. Base model NVIDIA quant ⚡ Turbo (this model) GPU memory 58.9 GiB 31 GiB 18.5 GiB ( 68% base, 40% nvidia) GPQA Diamond 75.71% 75.46% 72.73% ( 2.98% base, 2.73% nvidia) MMLU Pro 85.25% 84.94% 83.93% ( 1.32% base, 1.01% nvidia) Prefill 6352 tok/s 11069 tok/s 15359 tok/s (+142% base,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy