Language 中文 English Model Details This is a mixed precision W4A16 AWQ quantized version of Qwen/Qwen3 Coder 30B A3B Instruct, generated with llm compressor. Please follow the license of the original model. Quantization Strategy Layer Type Bits Notes Expert layers (128 experts) 4 bit MoE expert MLPs Non expert layers (attention, gate) 16 bit Higher precision for quality shared expert gate 16 bit Skipped (shape not divisible by 32) lm head 16 bit Skipped Model Size Bits Model Size Coding: LiveCodeBench v6 Multilingual: MMLU ProX OpenClaw: PinchBench Original BF16 ~60GB mixed W4A16 ~18GB ( 70% reduction ↓↓ ) 0.52 0.59 0.41 W4A16 ~17GB 0.51 0.57 0.37 Quickstart vLLM Usage vLLM is a high throughput and memory efficient inference and serving engine for LLMs. Directly talk to the model Directly use the OpenAPI See its documentation for more details. The following will create API endpoints at http://localhost:8000/v1 . Generate the Model Ethical Considerations and Limitations The model can produce factually incorrect output, and should not be relied on to produce factually accurate information. Because of the limitations of the pretrained model and the finetuning datasets, it is possible t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy