Llamacpp imatrix Quantizations of Qwen3 32B by Qwen Using llama.cpp release b5200 for quantization. Original model: https://huggingface.co/Qwen/Qwen3 32B All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Qwen3 32B bf16.gguf bf16 65.53GB true Full BF16 weights. Qwen3 32B Q8 0.gguf Q8 0 34.82GB false Extremely high quality, generally unneeded but max available quant. Qwen3 32B Q6 K L.gguf Q6 K L 27.26GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Qwen3 32B Q6 K.gguf Q6 K 26.88GB false Very high quality, near perfect, recommended . Qwen3 32B Q5 K L.gguf Q5 K L 23.69GB false Uses Q8 0 for embed and output weights. High quality, recommended . Qwen3 32B Q5 K M.gguf Q5 K M 23.21GB false High quality, recommended . Qwen3 32B Q5 K S.gguf Q5 K S 22.64GB false High quality, recommended . Qwen3 32B Q4 1.gguf Q4 1 20.64GB false Legacy format, similar performance to Q4 K S but with improved tokens/watt on Apple silicon. Qwen3 32B Q4 K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy