MiniMax M3 NVFP4 The first NVFP4 quantization of MiniMaxAI/MiniMax M3 — 428B total / 23B active MoE with MiniMax Sparse Attention (MSA), quantized 2026 06 12. ~256 GB on disk (vs 854 GB BF16, 444 GB MXFP8) → serves on 2× B300/GB300 with huge KV headroom, or fits tighter Blackwell pairs Routed + shared experts in NVFP4 (group 16, two level scaling); attention, MSA indexer, router, embeddings, lm head and the vision tower kept in original BF16 (no quantization round trip — copied verbatim from the source checkpoint) Same recipe family as nvidia/MiniMax M2.7 NVFP4 (experts only NVFP4), produced with TensorRT Model Optimizer 0.44.0 Quantization recipe Method PTQ, NVFP4 (FP4 weights+activations, FP8 per 16 block scales + FP32 global) Tool nvidia modelopt 0.44.0, transformers main (native minimax m3 vl ) Calibration 512 samples × 2048 tokens: cnn dailymail + nvidia/OpenCodeReasoning + nvidia/OpenMathReasoning Quantized routed experts (w1/w2/w3) + shared experts, all 57 MoE layers Excluded attention (incl. MSA indexer), router/gate, embeddings, lm head, vision tower, projectors KV cache not quantized (v1; MSA is young in engines — don't stack experiments) Calibration deliberately uses lon…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy