Gemma 4 26B A4B it NVFP4 First community NVFP4 quantization of google/gemma 4 26B A4B it — the Mixture of Experts variant of Gemma 4 with 25.2B total parameters and only 3.8B active per token. W4A4 — weights in FP4, activations in FP16 (full W4A4 quantization). Key Specs Original (BF16) NVFP4 (this) Size on disk ~49 GB ~16.5 GB Compression — 3.0x Total parameters 25.2B 25.2B Active parameters 3.8B 3.8B Architecture MoE: 128 experts, 8 active/token same Context window 256K tokens 256K tokens Modalities Text, Image, Video Text, Image, Video (all verified) Quantization — W4A4 (FP4 weights AND activations) Benchmarks A/B comparison against the BF16 original, both served via vLLM on DGX Spark (GB10 Blackwell, SM 12.1). Quality via lm evaluation harness with apply chat template . Quality Benchmark BF16 (reference) NVFP4 (this) Retained GSM8K (flexible extract) 87.79% 84.23% 95.9% GSM8K (strict match) 86.96% 82.64% 95.0% IFEval prompt strict 89.46% 87.99% 98.3% IFEval inst strict 92.81% 91.37% 98.4% IFEval prompt loose 90.94% 89.65% 98.6% IFEval inst loose 93.88% 93.05% 99.1% Average 90.31% 88.15% 97.6% Math reasoning (GSM8K) takes a ≈4pp hit — chained numerical steps accumulate rounding…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy