Qwen3 4B Instruct 2507 AWQ Method Quantised using vllm project/llm compressor, nvidia/Llama Nemotron Post Training Dataset and the following configs: Qwen3 4B Instruct 2507 Highlights We introduce the updated version of the Qwen3 4B non thinking mode , named Qwen3 4B Instruct 2507 , featuring the following key enhancements: Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage . Substantial gains in long tail knowledge coverage across multiple languages . Markedly better alignment with user preferences in subjective and open ended tasks , enabling more helpful responses and higher quality text generation. Enhanced capabilities in 256K long context understanding . Model Overview Qwen3 4B Instruct 2507 has the following features: Type: Causal Language Models Training Stage: Pretraining & Post training Number of Parameters: 4.0B Number of Paramaters (Non Embedding): 3.6B Number of Layers: 36 Number of Attention Heads (GQA): 32 for Q and 8 for KV Context Length: 262,144 natively . NOTE: This model supports only non thinking mode and does not generate blocks in its output. Mea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy