Version 26.05.01 Calibration STEM and Agentic Languages EN ZH HI AR RU JA KO NL FR ES Model Size 258.02 GB Contact Email Serving with vLLM This checkpoint needs a patched vLLM (MiniMax M3 compressed tensors support). The patch is Python only, so it installs on top of upstream's precompiled binaries — no CUDA compilation. Install Serve MiniMax M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters. Highlights: Native Multimodality: M3 undergoes mixed modality training from the very first step, enabling deeper semantic fusion across text, image, and video. Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per token compute to 1/20. Coding & Cowork Capability: M3 achieves frontier level performance across long horizon agentic benchmarks, excelling in both coding and cowork. MiniMax Sparse Attention (MSA) M3 is powered by MiniMax Sparse Attention (MSA) , a high performance sparse attention operator designed for million token contexts. Compared with GQA, MSA dramatically reduces the atte…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy