glm 4 9b chat GGUF Llama.cpp static quantization of THUDM/glm 4 9b chat Original Model: THUDM/glm 4 9b chat Original dtype: BF16 ( bfloat16 ) Quantized by: https://github.com/ggerganov/llama.cpp/pull/6999 IMatrix dataset: here Files Common Quants All Quants Downloading using huggingface cli Inference Simple chat template Chat template with system prompt Llama.cpp FAQ Why is the IMatrix not applied everywhere? How do I merge a split GGUF? Files Common Quants Filename Quant type File Size Status Uses IMatrix Is Split glm 4 9b chat.Q8 0.gguf Q8 0 9.99GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q6 K.gguf Q6 K 8.26GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q4 K.gguf Q4 K 6.25GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q3 K.gguf Q3 K 5.06GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q2 K.gguf Q2 K 3.99GB ✅ Available ⚪ Static 📦 No All Quants Filename Quant type File Size Status Uses IMatrix Is Split glm 4 9b chat.BF16.gguf BF16 18.81GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.FP16.gguf F16 18.81GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q8 0.gguf Q8 0 9.99GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q6 K.gguf Q6 K 8.26GB ✅ Available ⚪ Static 📦 No glm 4 9b chat.Q5 K.gguf Q5 K 7.14GB ✅ Ava…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy