Model Overview Model Architecture: Qwen3 5MoeForConditionalGeneration Input: Text, Image, Video Output: Text Supported Hardware Microarchitecture: AMD MI300 MI350/MI355 ROCm : 7.0.0 PyTorch : 2.9.1 Transformers : 5.3.0 Operating System(s): Linux Inference Engine: SGLang/vLLM Model Optimizer: AMD Quark (v0.12) Quantized layers: Experts in language model only Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic Model Quantization The model was quantized from Qwen/Qwen3.5 397B A17B FP8 using AMD Quark. The weights are quantized to MXFP4 and activations are quantized to MXFP4. Quantization scripts: For further details or issues, please refer to the AMD Quark documentation or contact the respective developers. Evaluation The model was evaluated on gsm8k benchmarks using the vllm framework. Accuracy Benchmark Qwen/Qwen3.5 397B A17B FP8 amd/Qwen3.5 397B A17B MXFP4(this model) Recovery gsm8k (flexible extract) 95.38 94.54 99.12% Reproduction The GSM8K results were obtained using the vLLM framework, based on the Docker image rocm/vllm dev:nightly main 20260211 , and vLLM is installed inside the container. Evaluating model in a new terminal License Modifications…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy