Model Overview Model Architecture: MiniMaxM3SparseForConditionalGeneration Input: Text, Image Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.1.1 PyTorch : 2.10.0 Transformers : 5.2.0 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic Model Quantization The model was quantized from MiniMaxAI/MiniMax M3 using AMD Quark. The weights are quantized to MXFP4 and activations are quantized to MXFP4. Quantization scripts: Evaluation The model was evaluated on gsm8k benchmarks using the vllm framework. Accuracy Benchmark MiniMaxAI/MiniMax M3 amd/MiniMax M3 MXFP4(this model) Recovery gsm8k (flexible extract) 95.30 94.19 98.84% Reproduction The GSM8K results were obtained using the lm eval framework, based on the Docker image rocm/pytorch private:vllm hy mm 06112026 . The vLLM shipped in that image was used as is, with the patch from this PR ( 45794) applied on top. Before running the evaluation, install the required packages: Launching server Evaluating model in a new terminal
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy