Qwen3 235B A22B Instruct 2507 FP8 Highlights We introduce the updated version of the Qwen3 235B A22B FP8 non thinking mode , named Qwen3 235B A22B Instruct 2507 FP8 , featuring the following key enhancements: Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage . Substantial gains in long tail knowledge coverage across multiple languages . Markedly better alignment with user preferences in subjective and open ended tasks , enabling more helpful responses and higher quality text generation. Enhanced capabilities in 256K long context understanding . Model Overview This repo contains the FP8 version of Qwen3 235B A22B Instruct 2507 , which has the following features: Type: Causal Language Models Training Stage: Pretraining & Post training Number of Parameters: 235B in total and 22B activated Number of Paramaters (Non Embedding): 234B Number of Layers: 94 Number of Attention Heads (GQA): 64 for Q and 4 for KV Number of Experts: 128 Number of Activated Experts: 8 Context Length: 262,144 natively . NOTE: This model supports only non thinking mode and does not generate blocks i…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy