Model Overview Model Architecture: Qwen3 5MoeForConditionalGeneration Input: Text, Image, Video Output: Text Supported Hardware Microarchitecture: AMD MI350 / MI355 ROCm : 7.2.0 PyTorch : 2.9.1 Transformers : 5.3.0 Operating System(s): Linux Inference Engine: SGLang Model Optimizer: AMD Quark (v0.12) Quantized layers: All MoE experts in the language model, including the shared expert (the shared expert is also fused into the MoE kernel for faster decode). Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic This checkpoint extends the routed expert MXFP4 quantization by also quantizing the shared expert to MXFP4 and fusing it into the routed MoE kernel (FSE: fused shared expert). Compared with keeping the shared expert in bf16, this further reduces the bf16 footprint and improves decode throughput, with no measurable accuracy loss on GSM8K (see Evaluation). Model Quantization The model was quantized from Qwen/Qwen3.5 397B A17B FP8 using AMD Quark. Weights and activations are quantized to OCP MXFP4. Quantization scripts: For further details or issues, please refer to the AMD Quark documentation or contact the respective developers. Evaluation The model…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy