MiniMax M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. Highlights: Native Multimodality: M3 undergoes mixed modality training from the very first step, enabling deeper semantic fusion across text, image, and video. Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per token compute to 1/20. Coding & Cowork Capability: M3 achieves frontier level performance across long horizon agentic benchmarks, excelling in both coding and cowork. MiniMax M3 MXFP8 is the MXFP8 quantized variant of MiniMax M3, a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. MiniMax Sparse Attention (MSA) M3 is powered by MiniMax Sparse Attention (MSA) , a high performance sparse attention operator designed for million token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality. 📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers How to Use MiniMax Agent MiniMax API M3…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy