GLM 5.2 REAP 504B GGUF — BF16 + dynamic 2/3/4 bit GGUF builds of the 34% expert pruned, Router KD recovered GLM 5.2 504B — for llama.cpp (CPU / Metal / CUDA). Includes the first working GGUF of a GLM 5.2 DSA model , made possible by a custom patch for its shared sparse attention indexer. The full precision NVFP4 model lives at 0xSero/GLM 5.2 504B — start there for the model story, eval, and architecture. This card is about the GGUF conversion specifically. 🙏 Sponsor Built on 8× NVIDIA B200 sponsored by Lambda . Thank you. 🙏 Files file bits size notes GLM 5.2 REAP 504B BF16 16 ~933 GB full precision — for fine tuning / re quantizing (23 shards) GLM 5.2 REAP 504B Q4 K XL ~4.5 ~325 GB recommended — dynamic, highest usable quality (8 shards) GLM 5.2 REAP 504B Q3 K XL ~3.5 ~259 GB dynamic, strong quality/size (6 shards) GLM 5.2 REAP 504B Q2 K XL ~2.7 ~111 GB dynamic, smallest (5 shards) Files 45 GB are split ( ... 00001 of 000NN.gguf ); point llama.cpp at the first shard and it loads the rest automatically. What "dynamic" means here These are not flat K quants. Each uses per tensor precision : the parts that hurt most under quantization are kept high while the bulk (routed experts) is…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy