Huihui gemma 4 31B it abliterated v2 MXFP4 MXFP4 quantized version of huihui ai/Huihui gemma 4 31B it abliterated v2 Quantization Details Property Value Date 2026 04 15 Scheme MXFP4A16 (4 bit weight only) Algorithm GPTQ Group Size 32 Output Size 18.2 GB Compression ~3.4x Format mxfp4 pack quantized (compressed tensors) Quantized Layers (~1,571 layers) Main LLM attention layers (q proj, k proj, v proj, o proj) MLP layers (gate proj, up proj, down proj) Embed tokens BF16 Layers (192 layers) Vision tower (26 encoder layers) Vision embeddings LM head Comparison: Original vs Quantized Metric Original (BF16) Quantized (MXFP4) Difference Size 59 GB 18.2 GB 69% (40.8 GB saved) Precision 16 bit 4 bit 4x reduction Quantization None mxfp4 pack Compressed tensors format Layers Quantized 0 1,571 All LLM layers Layers BF16 All 192 Vision tower only Memory Requirements Original (BF16) : ~62 GB VRAM for inference Quantized (MXFP4) : ~20 GB VRAM for inference (~3x less) Inference Speed (Estimated) With Blackwell GPU (GB10/B200) native FP4 tensor cores: Throughput : ~2 3x faster decoding vs BF16 Pre fill : Similar or slight improvement Quality (PPL) Model PPL Notes google/gemma 4 31B it (f16) 14,874…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy