Qwen3 235B A22B Thinking 2507 Highlights Over the past three months, we have continued to scale the thinking capability of Qwen3 235B A22B, improving both the quality and depth of reasoning. We are pleased to introduce Qwen3 235B A22B Thinking 2507 , featuring the following key enhancements: Significantly improved performance on reasoning tasks, including logical reasoning, mathematics, science, coding, and academic benchmarks that typically require human expertise — achieving state of the art results among open source thinking models . Markedly better general capabilities , such as instruction following, tool usage, text generation, and alignment with human preferences. Enhanced 256K long context understanding capabilities. NOTE : This version has an increased thinking length. We strongly recommend its use in highly complex reasoning tasks. Model Overview Qwen3 235B A22B Thinking 2507 has the following features: Type: Causal Language Models Training Stage: Pretraining & Post training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Numb…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy