The First Published NVFP4 Quantization of MiniMax M3 NOTICE This is an experimental quantization of MiniMax M3 to NVFP4 for use on DGX Spark. (Note: This quantization is not DGX Spark only.) For DGX Spark Users: Run with sparkrun; part of Spark Arena https://sparkrun.dev https://spark arena.com To run with sparkrun on 4x DGX Spark Nodes: For RTX Pro 6000 Users (or DGX Spark Users who don't want to use sparkrun): You can run this using the custom sglang container: (Container build is multi arch so it can be used for x86 and ARM) Reference settings can be derived from the sparkrun recipe: https://github.com/spark arena/recipe registry/blob/main/experimental recipes/minimax m3/minimax m3 v0 nvfp4 4x.yaml Happy Coding! Let's go! MiniMax M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. Highlights: Native Multimodality: M3 undergoes mixed modality training from the very first step, enabling deeper semantic fusion across text, image, and video. Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M contex…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy