Llamacpp imatrix Quantizations of Qwen2.5 14B Instruct 1M Using llama.cpp release b4546 for quantization. Original model: https://huggingface.co/Qwen/Qwen2.5 14B Instruct 1M All quants made using imatrix option with dataset from here Run them in LM Studio Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description Qwen2.5 14B Instruct 1M f32.gguf f32 59.09GB true Full F32 weights. Qwen2.5 14B Instruct 1M f16.gguf f16 29.55GB false Full F16 weights. Qwen2.5 14B Instruct 1M Q8 0.gguf Q8 0 15.70GB false Extremely high quality, generally unneeded but max available quant. Qwen2.5 14B Instruct 1M Q6 K L.gguf Q6 K L 12.50GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . Qwen2.5 14B Instruct 1M Q6 K.gguf Q6 K 12.12GB false Very high quality, near perfect, recommended . Qwen2.5 14B Instruct 1M Q5 K L.gguf Q5 K L 10.99GB false Uses Q8 0 for embed and output weights. High quality, recommended . Qwen2.5 14B Instruct 1M Q5 K M.gguf Q5 K M 10.51GB false High quality, recommended . Qwen2.5 14B Instruct 1M Q5 K S.gguf Q5 K S 10.27GB false High quality, recommended . Qwen2.5 14B Instruct 1M Q4 K L.ggu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy