EXPERIMENTAL AWQ quantization: done by stelterlab in INT4 GEMM with llm compressor (https://github.com/vllm project/llm compressor) from the vllm project. Still trying to get familiar with llm compressor. Tested against vLLM v0.9.2. Original Weights by Qwen AI. Original Model Card follows: Qwen3 30B A3B Instruct 2507 Highlights We introduce the updated version of the Qwen3 30B A3B non thinking mode , named Qwen3 30B A3B Instruct 2507 , featuring the following key enhancements: Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage . Substantial gains in long tail knowledge coverage across multiple languages . Markedly better alignment with user preferences in subjective and open ended tasks , enabling more helpful responses and higher quality text generation. Enhanced capabilities in 256K long context understanding . Model Overview Qwen3 30B A3B Instruct 2507 has the following features: Type: Causal Language Models Training Stage: Pretraining & Post training Number of Parameters: 30.5B in total and 3.3B activated Number of Paramaters (Non Embedding): 29.9B Number of Layers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy