Intro The AWQ version is quantized using ms swift. You may refer to our best practice for training/fine tuning Qwen3 models here. Note that the AWQ version for Qwen3 MoE models are verified to be working on Transformers/vLLM. We have not have the chance to tested them on other engines. Inference Quantization The model has undergone AWQ int4 quantization using the ms swift framework. Since the model is based on the MoE (Mixture of Experts) architecture, all linear layers except for gate and lm head have been quantized. If you have fine tuned the model and wish to quantize the fine tuned version, you can refer to the following quantization scripts: Dense Model Quantization Script: View Here MoE Model Quantization Script: View Here With these scripts, you can easily complete the quantization process for the model. Evaluation We evaluate the quality of this AWQ quantization with EvalScope. For the best practice for evaluating Qwen3 models, one may refer to the following: 最佳实践 Best Practice Performance of Qwen3 30B A3B AWQ is evaluated on our mixed benchmark of Qwen3 Evaluation Collection, with the results listed below: The performance comparison of Qwen3 30B A3B AWQ and Qwen3 30B A3B t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy