Qwen3 Next 80B A3B Instruct quantized.w4a16 Model Overview Model Architecture: Qwen3NextForCausalLM Input: Text Output: Text Model Optimizations: Weight quantization: INT4 Version: 1.0 Model Developers: RedHat (Neural Magic) ModelCar Storage URI : oci://registry.redhat.io/rhai/modelcar qwen3 next 80b a3b instruct quantized w4a16:3.0 Validated on vLLM: 0.13.0 Validated on RHAIIS: 3.3 Validated on RHOAI: 3.3 Model Optimizations This model was obtained by quantizing the weights of Qwen/Qwen3 Next 80B A3B Instruct to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. Only the weights of the linear operators within transformers blocks are quantized. Weights are quantized using a symmetric per group scheme, with group size 128. The GPTQ algorithm is applied for quantization, as implemented in the llm compressor library. Deployment This model can be deployed efficiently using the vLLM backend, as shown in the example below. vLLM aslo supports OpenAI compatible serving. See the documentation for more details. Creation Creation details This model was created with llm compressor by…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy