Qwen3 32B GPTQ Int4 Base model: Qwen/Qwen3 32B This model is quantized to 4 bit with a group size of 128. Compared to earlier quantized versions, the new quantized model demonstrates better tokens/s efficiency. This improvement comes from setting desc act=False in the quantization configuration. 【Dependencies】 【Model Download】 【Overview】 Qwen3 32B Qwen3 Highlights Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture of experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction following, agent capabilities, and multilingual support, with the following key features: Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non thinking mode (for efficient, general purpose dialogue) within single model , ensuring optimal performance across various scenarios. Significantly enhancement in its reasoning capabilities , surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non thinking mode) on mathematics, code generation, and commonsense logical reasoning. Superior human prefere…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy