MiniMax M3 FP8 dynamic Model Overview This model is an FP8 dynamic quantized version of MiniMaxAI/MiniMax M3. Base model: MiniMaxAI/MiniMax M3 Optimization: FP8 dynamic quantization Format: safetensors / compressed tensors Validated runtime: vLLM OpenAI compatible server Tested hardware: AMD MI350, tensor parallel size 8 MiniMax M3 is a native multimodal MoE model. The original model card describes it as a ~428B parameter model with ~23B activated parameters and 1M context support. License This quantized checkpoint follows the license terms of the base model, MiniMaxAI/MiniMax M3. The Hugging Face model card metadata uses license: other because the MiniMax community license is not one of the Hub's enumerated license identifiers. Model Optimizations This checkpoint uses FP8 dynamic quantization to reduce memory and disk requirements while preserving model quality. Validation below compares this quantized checkpoint against the BF16 MiniMaxAI/MiniMax M3 baseline. Evaluation The model was evaluated against BF16 MiniMaxAI/MiniMax M3 . Scores are averaged across seeds. Benchmark MiniMaxAI/MiniMax M3 EmbeddedLLM/MiniMax M3 FP8 dynamic Recovery (%) : : : GSM8k Platinum 95.81 95.92 100.12…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy