GLM 5.2 GGUF 2.244bpw This is a 2.2 BPW quantized model for the GPU riches with more combined RAM + VRAM than common sense. The quant aims to achieve best in class performance, by relying on SOTA quants from ik llama.cpp: Routed experts tensors use the IQ2 KT quant (2.125 BPW) Indexer tensors use either the Q6 0 quant (6.5 BPW) or the Q8 0 quant (8.5 BPW) All other tensors use the Q6 0 quant Coupled with the recent enhancements: MTP support with e.g. spec type mtp:n max=4,p min=0.0 ( 1890) graph parallel support with sm graph ( 1821) DSA support with dsa fidx ( 2045, 2098, 2109, and many others) it should run at decent speed as well, with very little slowdown at long context. (Note: For now, sm graph and quantize KV cache e.g. ctk q8 0 do not work together with dsa fidx . As a fun exercise, you can ask this quant to get them to work.) Versions There are 3 versions: GLM 5.2 GGUF 2.244bpw.gguf Made with the imatrix from unsloth (thanks!) GLM 5.2 GGUF 2.244bpw q8indexer.gguf Same as the above, but with Q8 0 for the indexer tensors GLM 5.2 GGUF 2.244bpw muzzy imatrix.gguf Same as the above, but with the imatrix from muzzy (thanks!) Comparison: version imatrix indexer ppl GLM 5.2 GGUF 2…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy