Llamacpp imatrix Quantizations of Qwen3 1.7B by Qwen Using llama.cpp release b5200 for quantization. Original model: https://huggingface.co/Qwen/Qwen3 1.7B All quants made using imatrix option with dataset from here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Qwen3 1.7B bf16.gguf bf16 4.07GB false Full BF16 weights. Qwen3 1.7B Q8 0.gguf Q8 0 2.17GB false Extremely high quality, generally unneeded but max available quant. Qwen3 1.7B Q6 K L.gguf Q6 K L 1.82GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Qwen3 1.7B Q6 K.gguf Q6 K 1.67GB false Very high quality, near perfect, recommended . Qwen3 1.7B Q5 K L.gguf Q5 K L 1.66GB false Uses Q8 0 for embed and output weights. High quality, recommended . Qwen3 1.7B Q4 K L.gguf Q4 K L 1.51GB false Uses Q8 0 for embed and output weights. Good quality, recommended . Qwen3 1.7B Q5 K M.gguf Q5 K M 1.47GB false High quality, recommended . Qwen3 1.7B Q5 K S.gguf Q5 K S 1.44GB false High quality, recommended . Qwen3 1.7B Q3 K XL.gguf Q3 K XL 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy