Model Overview Model Architecture: qwen3 next Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm: 7.1.0 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark (V0.11) moe Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic attn: linear attn.out proj , self attn.o proj Weight quantization: OCP MXFP4, Static Activation quantization: OCP MXFP4, Dynamic Calibration Dataset: Pile This model was built with Qwen3 Coder Next model by applying AMD Quark for MXFP4 quantization. Model Quantization The model was quantized from Qwen/Qwen3 Coder Next using AMD Quark. The weights and activations are quantized to MXFP4. Quantization scripts: Note that qwen3 next is not in the built in model template list in Quark V0.11, it has to be registered before quantization. Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Evaluation The model was evaluated on GSM8K benchmarks. Accuracy Benchmark Qwen3 Coder Next Qwen3 Coder Next MXFP4(this model) Recovery GSM8K (flexible extract) 94.54 93.25 98.6% Reproduction The GSM8K results were obtained using the lm evaluation harness framework,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy