Qwen3 4B Instruct 2507 GPTQ Int4 Languages: Multilingual support (104 languages, including languages of the CIS countries). Model Description This is a quantized version of Qwen/Qwen3 4B Instruct 2507. The model was quantized using llmcompressor with GPTQ method (W4A16). Quantization: GPTQ 4 bit (weights), 16 bit (activations) Format: compressed tensors (native vLLM support) Group Size: 128 Act Order: Static How to Run vLLM (Recommended) This format is optimized for vLLM. Python (using vLLM) License Apache 2.0 (Same as original model).
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy